Text similarity intelligent analysis method

By using the Resemblyzer model to process text data in real time, convert it into speech signals and compare vectors, the problems of weak semantic understanding, poor context sensitivity and low computational efficiency of traditional text similarity analysis methods are solved, and efficient and accurate text similarity analysis is achieved.

CN120235136APending Publication Date: 2025-07-01BEIYIN FINANCIAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411636437.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Traditional text similarity analysis methods have problems such as weak semantic understanding, poor context sensitivity and low computational efficiency, making it difficult to achieve efficient and accurate text similarity analysis in a big data environment.

Method used

The preprocessed text data is processed in real time by using the Resemblyzer model, and the text-to-speech TTS technology is used to convert the text into a speech signal, and the voice encoder is used to convert the speech signal into a high-dimensional vector representation, and vector comparison is performed to evaluate the text similarity.

Benefits of technology

It realizes efficient text similarity analysis, improves computing efficiency and accuracy, can run stably in a big data environment, and generates detailed text similarity analysis reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235136A_ABST
    Figure CN120235136A_ABST
Patent Text Reader

Abstract

The invention discloses a text similarity intelligent analysis method which comprises the following steps: receiving text data uploaded by a user, and preprocessing the text data to obtain preprocessed data; performing real-time processing on the pre-processed data by adopting a Resemblyzer model to obtain a processing result; and generating a text similarity analysis report according to the processing result. The system not only can realize efficient text similarity analysis, but also can ensure stable performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text similarity analysis, and in particular to an intelligent analysis method for text similarity. Background Art

[0002] Currently, in the big data environment, text data processing and analysis have become an important direction for technological development. With the popularization of the network and the development of technology, the amount of text data is increasing continuously. However, traditional text similarity analysis methods face problems such as low computational efficiency and susceptibility to noise. The Resemblyzer model can convert text into a vector form for representation, which is suitable for the comparison of large-scale text data.

[0003] Existing text similarity calculation methods mainly rely on traditional statistical methods such as the bag-of-words model or TF-IDF. They represent text as a word frequency vector and can calculate similarity through methods such as Euclidean distance.

[0004] Disadvantages of the prior art:

[0005] Weak semantic understanding: The traditional bag-of-words model cannot capture the semantic relationships between words in the text, making it difficult to accurately locate the core content of the text and resulting in inaccurate similarity calculation results.

[0006] Poor context sensitivity: Traditional text similarity calculation methods usually ignore the context information between words in the text, so it is difficult to accurately reflect the degree of semantic similarity, especially for long texts or texts with nested structures.

[0007] Low computational efficiency: When dealing with large text data sets, traditional calculation methods will face high computational complexity, affecting the real-time performance and application scope of the system. Summary of the Invention

[0008] In view of the above problems, the present invention is proposed to provide an intelligent analysis method for text similarity that overcomes the above problems or at least partially solves the above problems.

[0009] According to one aspect of the present invention, an intelligent analysis method for text similarity is provided. The intelligent analysis method includes:

[0010] Receiving text data uploaded by a user and performing preprocessing to obtain preprocessed data;

[0011] Performing real-time processing on the preprocessed data using the Resemblyzer model to obtain a processing result;

[0012] Generating a text similarity analysis report from the processing result.

[0013] Optionally, the receiving of the text data uploaded by the user includes:

[0014] Receive the text data uploaded by the user, supporting text files in multiple formats or the text content input in real time.

[0015] Optionally, the preprocessing to obtain the preprocessed data specifically includes:

[0016] Perform cleaning, word segmentation, and stop word removal preprocessing operations on the input text data, and unify the format to prepare for subsequent text conversion.

[0017] Optionally, the real-time processing of the preprocessed data by using the Resemblyzer model to obtain the processing result specifically includes:

[0018] Convert the preprocessed text into a voice signal;

[0019] Perform a vector comparison engine.

[0020] Optionally, the conversion of the preprocessed text into a voice signal specifically includes: using text-to-speech (TTS) technology to convert the preprocessed text into a voice signal.

[0021] Optionally, the performing of the vector comparison engine specifically includes:

[0022] Use the voice encoder in the Resemblyzer model to convert the converted voice signal into a high-dimensional vector representation;

[0023] Evaluate the similarity between texts by comparing the similarity between the vectors corresponding to different texts.

[0024] Optionally, the generating of the text similarity analysis report from the processing result specifically includes:

[0025] Generate a text similarity analysis report according to the result of the vector comparison, and display the result to the user.

[0026] According to a method for intelligent analysis of text similarity as claimed in claim 7, wherein the text similarity analysis report includes a similarity score and key difference point information.

[0027] A method for intelligent analysis of text similarity provided by the present invention, the intelligent analysis method comprising: receiving the text data uploaded by the user, and performing preprocessing to obtain preprocessed data; performing real-time processing on the preprocessed data by using the Resemblyzer model to obtain a processing result; generating a text similarity analysis report from the processing result. A system that can achieve efficient text similarity analysis and ensure stable performance.

[0028] The above description is only an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention, it can be implemented according to the content of the specification. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Brief Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0030] Figure 1 It is a flowchart of a method for intelligent analysis of text similarity provided by an embodiment of the present invention;

[0031] Figure 2 It is a block diagram of the composition of a system corresponding to a method for intelligent analysis of text similarity provided by an embodiment of the present invention. Detailed Description of the Embodiments

[0032] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0033] The terms "including" and "having" and any variations thereof in the embodiments of the specification, claims and drawings of the present invention are intended to cover non-exclusive inclusion. For example, including a series of steps or units.

[0034] The technical solutions of the present invention will be further described in detail below with reference to the drawings and embodiments.

[0035] As Figure 1 shown, a method for intelligent analysis of text similarity, the intelligent analysis method includes: receiving text data uploaded by a user and performing preprocessing to obtain preprocessed data; using the Resemblyzer model to perform real-time processing on the preprocessed data to obtain a processing result; generating a text similarity analysis report based on the processing result.

[0036] As Figure 2 shown, a text similarity intelligent analysis system based on the Resemblyzer model.

[0037] The system includes: a data acquisition module, a Resemblyzer model implementation module, a result generation and output module, and a user interface.

[0038] Data acquisition module:

[0039] The data input module is responsible for receiving the text data uploaded by the user and supports text files in multiple formats or text content input in real time.

[0040] Preprocessing: Perform preprocessing operations such as cleaning, word segmentation, and stop word removal on the input text data to unify the format for subsequent text conversion.

[0041] Resemblyzer processing module, including:

[0042] As the core processing unit, the Resemblyzer model is used to perform real-time processing on the text data uploaded by the user. The specific steps include:

[0043] Text conversion: Use text-to-speech (TTS) technology to convert the preprocessed text into a speech signal. This step is one of the innovations of the present invention. By converting text into speech, the advantages of the Resemblyzer model in speech processing are utilized to analyze text similarity.

[0044] Vector comparison engine: Use the voice encoder in the Resemblyzer model to convert the converted speech signal into a high-dimensional vector representation (embedding). Then, by comparing the similarity between the vectors corresponding to different texts, the similarity degree between texts is evaluated.

[0045] Result generation and output module, including:

[0046] According to the result of vector comparison, generate a text similarity analysis report, including information such as similarity scores and key difference points, and display the result to the user through the output control module.

[0047] User interface, including:

[0048] Provide a friendly user interface that allows users to upload text files, input real-time text, view the similarity analysis results, and support the export and sharing of the results.

[0049] Model overview

[0050] The Resemblyzer model is a pre-trained language model based on the Transformer architecture, with powerful semantic understanding capabilities and sensitivity to context. By learning a large amount of text data, Resemblyzer can capture the deep semantic relationships between words and generate more accurate text representations, providing a more effective method for text similarity calculation.

[0051] Resemblyzer is a text similarity analysis method based on the pre-trained language model Transformer. It utilizes the powerful semantic understanding ability of the Transformer model to effectively capture long-distance dependencies and context information in text, thereby improving the calculation accuracy of text similarity. Resemblyzer has learned rich language knowledge during the pre-training stage and can perform text similarity analysis without additional labeled data. At the same time, the Resemblyzer model is relatively small and has a low computational complexity, enabling efficient operation on devices with limited resources.

[0052] Application Areas:

[0053] The Resemblyzer model is widely used in multiple fields, including:

[0054] Security Verification: Used in security systems for voiceprint recognition, such as unlocking in telephone banking or smart home devices.

[0055] Media Production: Audio editing and mixing to generate realistic dialogue scenes for movies and games.

[0056] Educational Tools: Voice teaching and language learning applications that can provide personalized feedback.

[0057] Entertainment Industry: Music style conversion to produce unique music works.

[0058] Artificial Intelligence Assistants: Improve the interaction experience of virtual assistants to enable them to better understand and respond to users' voice commands.

[0059] Fine-tuning: In addition to the pre-trained model, Resemblyzer also allows users to fine-tune on specific datasets to adapt to specific application scenarios.

[0060] Pairwise similarity: Provides an effective method to calculate the similarity between two audio segments, which is very useful for scenarios such as personalized recommendations or audio classification.

[0061] One-shot learning: Supports identifying whether a new audio matches it through one example, which has wide applications in fields such as voiceprint recognition.

[0062] Voice cloning: It can perform simple audio style conversion to create new audio similar to the source audio.

[0063] Beneficial effects: Provide a system that can achieve both efficient text similarity analysis and ensure stable performance. Ways to achieve the goal: By optimizing the running efficiency of the Resemb l yzer model and using high-performance hardware and algorithm optimization strategies.

[0064] The above specific implementation manners have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only the specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A text similarity intelligent analysis method, characterized in that: The intelligent analysis method comprises: Receive text data uploaded by users, and perform preprocessing to obtain preprocessed data; Using a Resemblyzer model to process the preprocessed data in real time to obtain a processing result; The processing results are used to generate a text similarity analysis report.

2. According to claim 1, a text similarity intelligent analysis method is characterized in that: The receiving of text data uploaded by the user comprises: Receive text data uploaded by users, support text files in multiple formats or text content input in real time.

3. A text similarity intelligent analysis method according to claim 1, characterized in that: The preprocessing to obtain preprocessed data specifically includes: The input text data is preprocessed by cleaning, word segmentation, and stop word removal, and the format is unified to prepare for subsequent text conversion.

4. A text similarity intelligent analysis method according to claim 1, characterized in that: The use of the Resemblyzer model to process the preprocessed data in real time to obtain the processing results specifically includes: Convert the preprocessed text into speech signals; Performs a vector comparison engine.

5. A text similarity intelligent analysis method according to claim 4, characterized in that: The converting the preprocessed text into a speech signal specifically includes: using a text-to-speech (TTS) technology to convert the preprocessed text into a speech signal.

6. A text similarity intelligent analysis method according to claim 4, characterized in that: The vector comparison engine specifically includes: Using the sound encoder in the Resemblyzer model, the converted speech signal is converted into a high-dimensional vector representation; The similarity between texts is evaluated by comparing the similarities between vectors corresponding to different texts.

7. The method for intelligent text similarity analysis according to claim 1, characterized in that: The generating of the text similarity analysis report from the processing result specifically includes: Based on the results of the vector comparison, a text similarity analysis report is generated and the results are displayed to the user.

8. A text similarity intelligent analysis method according to claim 7, characterized in that: The text similarity analysis report includes similarity scores and key difference point information.