Cboth case detection and identification method and system, electronic equipment and medium
By using a multimodal content processing engine and a multi-model collaborative detection mechanism, the problem of narrow detection range and low accuracy in existing copywriting detection technologies has been solved, achieving efficient and accurate detection of multimodal content and adapting to business changes and rule updates.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-03-17
AI Technical Summary
Existing copywriting detection technologies rely excessively on preset rules and training corpora in specific domains, resulting in a narrow detection range, low detection accuracy, inability to respond promptly to business changes and new rules, and inability to effectively handle multimodal content.
Employing a multimodal content processing engine, a dynamic knowledge base update mechanism, and a multi-model collaborative detection mechanism, the system achieves automated detection of multimodal content through the collaborative work of the access layer, business service layer, and model layer. It utilizes at least two detection models for multi-level detection, generating a structured set of questions and a set of modification suggestions.
It improves the detection range and accuracy of copywriting detection, reduces the reliance on preset rules and domain training corpora, enhances the system's response speed and generalization ability, supports multiple access methods, and adapts to business changes and new rules.
Smart Images

Figure CN121683803A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of text detection, and more particularly to a script detection and identification method and system, an electronic device and a storage medium. BACKGROUND
[0002] With the explosive growth of the digital content industry, the efficiency of text creation and dissemination has been significantly improved, but the problems of originality protection, quality control and compliance review have become prominent. Traditional manual review is inefficient and standards are not unified, and AI technology breakthrough provides a new path for text automation processing. In response to this, script detection technology has emerged, which automatically identifies text grammar errors and other problems, and has become a key tool for content quality assurance.
[0003] However, the existing common script detection technology relies too much on preset rules and specific domain training corpus, which makes the system slow to respond when facing business changes or new rules, and weak in generalization ability, thereby causing the problems of narrow detection range and low detection accuracy.
[0004] Therefore, how to solve the problems of narrow detection range and low detection accuracy of traditional script detection technology is a technical problem that needs to be solved by those skilled in the art. SUMMARY
[0005] The present application provides a script detection and identification method, system, electronic device and medium, which has the characteristics of wide detection range and high detection accuracy.
[0006] To achieve the above-mentioned purpose, the first aspect of the present application provides a script detection and identification method, which comprises: receiving a script detection request sent by a business system, the script detection request comprising multi-modal content; performing text extraction processing on each modality data in the multi-modal content based on the text extraction algorithm corresponding to each modality, to obtain a to-be-detected text; detecting the to-be-detected text by at least two detection models to generate a detection result, the detection result comprising a structured problem set and a modification suggestion set; sending the detection result to the business system.
[0007] To achieve the above-mentioned purpose, the second aspect of the present application provides a script detection and identification system, which comprises an access layer, a business service layer and a model layer; The access layer is used to receive a script detection request sent by a business system, the script detection request comprising multi-modal content; The model layer is used to perform text extraction processing on each modality data in the multi-modal content based on the text extraction algorithm corresponding to each modality, to obtain a to-be-detected text; and to detect the to-be-detected text by at least two detection models; The business service layer is configured to generate a detection result including a structured question set and a modification suggestion set based on the detection model. The access layer is further configured to send the detection result to the business system.
[0008] To achieve the above object, the third aspect of the present application provides a text detection and recognition device, which comprises: A receiving unit is configured to receive a text detection request sent by a business system, wherein the text detection request comprises multi-modal content. An analysis unit is configured to perform text extraction processing on each modality data in the multi-modal content based on a text extraction algorithm corresponding to each modality, to obtain a to-be-detected text. A detection unit is configured to detect the to-be-detected text by using at least two detection models, to generate a detection result including a structured question set and a modification suggestion set. A sending unit is configured to send the detection result to the business system.
[0009] To achieve the above object, the fourth aspect of the present application provides an electronic device, which comprises: A memory is configured to store a computer program. A processor is configured to execute the computer program to implement the steps of the above-mentioned text detection and recognition method.
[0010] To achieve the above object, the fifth aspect of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the above-mentioned text detection and recognition method.
[0011] To achieve the above object, the sixth aspect of the present application provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of the above-mentioned text detection and recognition method.
[0012] Therefore, the present application has the following advantages: The application provides a text detection and identification method, which first receives a text detection request sent by a business system, wherein the text detection request comprises multi-modal content; text extraction processing is performed on each modal data in the multi-modal content based on a text extraction algorithm corresponding to each modal to obtain to-be-detected text; the to-be-detected text is detected by at least two detection models to generate a detection result, wherein the detection result comprises a structured question set and a modification suggestion set; and the detection result is sent to the business system. In this process, after receiving the detection request covering the multi-modal content, the text extraction algorithm corresponding to each modal can automatically extract text information from unstructured documents such as pictures and PDFs, and uniformly convert the multi-modal content into detectable text information. In addition, the system uses at least two detection models for detection, which can fully exert the advantages of different models and then identify various types of errors. The finally generated detection result comprises a structured question set and a modification suggestion set, which provides a high-quality data basis for subsequent model optimization.
[0013] Through the above multi-dimensional optimization, the excessive dependence on preset rules and domain training corpus can be effectively reduced, the response speed, generalization ability and detection precision can be improved, the detection range and accuracy can be improved, and the business changes and new rules can be adapted.
[0014] The application also discloses an electronic device, a computer readable storage medium and a computer program product, which can also achieve the above technical effects.
[0015] It should be understood that the above general description and the following detailed description are only exemplary and cannot limit the application. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or the prior art description. Obviously, the drawings in the following description only constitute some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor. The drawings are used to provide further understanding of the disclosure and constitute a part of the specification, and are used to explain the disclosure together with the following specific embodiments, but do not constitute a limitation on the disclosure. In the drawings: Figure 1 A structural schematic diagram of a text detection and identification system provided by an embodiment of the application; Figure 2 A flowchart of an embodiment of a text detection and identification method provided by an embodiment of the application; Figure 3 A flowchart of another embodiment of a text detection and identification method provided by an embodiment of the application; Figure 4 A text type data detection result schematic diagram provided by an embodiment of the present application is shown in FIG. 1. Figure 5 A structured configuration type data detection result schematic diagram provided by an embodiment of the present application is shown in FIG. 2. Figure 6 A picture type data and file type data detection result schematic diagram provided by an embodiment of the present application is shown in FIG. 3. Figure 7 A structure schematic diagram of a text detection and recognition device provided by an embodiment of the present application is shown in FIG. 4. Figure 8 A structure diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 5. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0018] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards in relevant regions.
[0019] In order to better understand the embodiments of the present application, first introduce the technical terms appearing in the present application: (1) CI / CD: Continuous Integration / Continuous Delivery: A software development practice that makes software delivery more frequent and reliable through automated build, test, and deployment processes. CI / CD pipeline can automatically detect code changes, run tests, and deploy code that passes tests to production, thereby accelerating software development cycle and improving quality.
[0020] (2) IDE: Integrated Development Environment: An application that provides a comprehensive set of tools for software development, usually including code editor, compiler, debugger and graphical user interface building tools. IDE aims to maximize programmer productivity, providing an integrated development experience.
[0021] (3) VSCode: Visual Studio Code: A lightweight yet powerful source code editor developed by Microsoft, supporting multiple programming languages, providing features such as intelligent code completion, syntax highlighting, code refactoring, debugging, and more. VSCode supports a rich extension ecosystem, allowing developers to customize their development environment as needed.
[0022] (4) TCQ: Testing, Compliance & Quality Team: A professional team responsible for ensuring that software products meet quality standards, comply with regulatory requirements, and pass comprehensive testing. The TCQ team is usually responsible for developing quality standards, implementing testing plans, monitoring compliance, identifying quality issues, and promoting continuous improvement. In the context of the text detection system, the TCQ team is responsible for evaluating the accuracy of the detection results and providing feedback for system optimization.
[0023] Currently, in the field of content and text quality control, text detection technology has become a core tool for enterprise compliance operation and risk prevention. Common text detection technologies on the market include the following: (1) Rule-based text correction system: Detect basic problems such as spelling errors and misuse of punctuation through pre-set grammar rules and dictionary libraries, such as Microsoft Word's spelling check function, which usually relies on manually written rule libraries and mainly detects spelling and basic grammar errors.
[0024] (2) Statistical language model detection: Use N-gram and other statistical models to detect unusual word combinations or expressions in text and determine possible errors.
[0025] (3) Traditional machine learning text classification method: Use SVM, random forest, and other algorithms to perform binary classification (such as compliance / non-compliance) on text. Generally, feature engineering is used to extract word frequency, syntactic structure, and other indicators, and is suitable for batch content screening.
[0026] (4) Single-modal AI detection system: Only supports pure text input and cannot process text information in multi-modal content such as images and PDFs. For example, the detection tool can analyze document text, but cannot identify illegal slogans in images.
[0027] (5) Independent deployment of content review system: Runs as a standalone application and requires manual import and export of content to be detected. For example, enterprises need to copy user-uploaded text into the review system, and manually return the results after detection is complete.
[0028] It should be noted that the text detection techniques (including but not limited to rule matching, N-gram statistics, SVM classification, etc.) listed in the embodiments of the present application are only illustrative examples, and any improved scheme based on the principles of the present application (such as a detection model introducing an attention mechanism, a semantic analysis combined with a knowledge graph, etc.) falls within the scope of protection of the present patent.
[0029] However, the above-mentioned example text detection techniques still have significant defects and significant limitations in dealing with complex business scenarios. The applicant found that due to the excessive dependence of existing text detection techniques on preset rules and specific domain training corpus, the following problems occur: (1) Detection capability limitation. Since traditional systems rely on manually defined rules, rule generation requires demand analysis, rule writing, testing and verification, etc., resulting in a long average time period. In addition, rule library updates require manual review, and the iteration speed of business rules far exceeds the efficiency of rule updates, resulting in rule update lag. In rule-based text correction systems, rule matching is mainly based on dictionary + regular, if-else, string matching, etc. Assuming that in a music-related text detection scenario, the manually defined rule dictionary has the word "luxury green diamond". When the text contains the expression "open luxury green diamond and enjoy high-quality music", the rule-based text correction system will find that "luxury" is not in the legal word library through regular matching, but since it is only based on the preset rules, it can only detect that there is a word out of the word library, but it is difficult to understand that "luxury" is semantically incorrect in this context, and it is even more difficult to accurately predict that it should be "luxury".
[0030] And in the detection based on statistical language model and traditional machine learning method, although the understanding model can understand that "luxury" is semantically incorrect in the context, but in the scenario where the text is about the album of "singer 1" but mistakenly mentions "song 1" of "singer 2", since the training data only includes "singer 2 - song 1", i.e. cannot accurately identify the album artist error, therefore this kind of statistical language model detection method depends on training data and rule library, if the training data and rule library are updated laggingly for the definition rules of this kind of artist name and work error collocation, the inference model in the traditional machine learning method may also be unable to accurately determine that it is a violation of the expression, thus reflecting the problem of limited detection accuracy due to rule update lag.
[0031] For this reason, due to the excessive dependence of existing text detection techniques on preset rules, the rule library presents a static feature, with an update cycle of several weeks to several months, which cannot respond to the dynamic changes of business rules in a timely manner, making it difficult to identify complex semantic errors, business logic errors and violation expressions, and the detection accuracy is limited.
[0032] (2) Modal support is single, and most systems only support pure text detection, and cannot directly process multi-modal content containing text such as pictures and PDFs. Assuming that a music album promotion material is presented in PDF format, which contains album poster pictures (pictures with album name, artist name, etc. Text information), song list, and some promotional texts, etc. Many existing text detection systems, due to single modal support, can only detect the pure text part (such as promotional text content). For the text information on the picture, the system cannot directly identify and process it, and cannot detect whether the text on the picture has incorrect expressions, illegal content, etc. Similarly, for some special layout or picture combined text content in PDF, due to the inability to effectively analyze multi-modal information, the error or illegal content in these parts may be missed, resulting in inaccurate overall detection results and a significant increase in multi-modal content missing rate.
[0033] (3) The process of independent deployment is fragmented, and the independent deployment system needs to manually import and export content. For example, after the music editor of a music platform completes the singer's promotional text, it needs to be exported from the CMS and then imported into the detection system for detection, and the result is presented in a specific report. Due to the lack of standard API interface, it cannot be automatically synchronized to business systems such as CRM and CMS. Integrating it with the CI / CD tool chain requires a lot of custom development, and without adaptive interface logic, the development team takes weeks or even months to write adaptive code. Even if it is integrated, the detection results cannot be synchronized in real time, and the audit personnel need to manually copy the results to the business system for modification.
[0034] This is because the existing technology relies too much on specific domain corpus, the system architecture is independent, lacks compatibility, open interface and standard data interaction, has weak integration capability, interrupts the audit process, prolongs the processing time, and affects platform content update and user experience.
[0035] To this end, in the embodiments of the present application, a multi-modal content processing engine, a knowledge base dynamic updating mechanism, and a multi-model collaborative detection mechanism are used to effectively improve the detection range and detection accuracy; and support multiple access methods, which can be flexibly integrated into different business scenarios. To facilitate understanding of the specific implementation of the script detection and recognition method provided by the embodiments of the present application, the following will be described with reference to the accompanying drawings.
[0036] It should be noted that the subject implementing the script detection and recognition method can be a script detection and recognition system provided by the embodiments of the present application, or a script detection and recognition apparatus provided by the embodiments of the present application. The script detection and recognition apparatus can be carried in an electronic device or a functional module of an electronic device. The electronic device in the embodiments of the present application can be any device capable of implementing the script detection and recognition method in the embodiments of the present application, for example, an Internet of Things (IoT) device.
[0037] In order to make the method provided by the embodiments of the present application clearer and easier to understand, first, the document detection and recognition system in the embodiments of the present application is introduced. The system in the embodiments of the present application can be seen from, for example, the document detection and recognition system 100 shown in Figure 1 The system can include, for example, an access layer 101, a service layer 102, a model layer 103 (Dify Flow and model provider), and an independent service 104.
[0038] (1) Access layer 101: As the entrance channel of the document detection and recognition system and external business systems, it undertakes two core functions of receiving requests / data and outputting data. Among them, the access layer 101 can receive the document detection request sent by the business system through the API gateway (processing HTTP request, providing authentication, flow limiting), CLI interface (supporting local call / CI / CD integration), web interface (providing user interaction interface) and other interfaces, and send the detection result (including structured question set, modification suggestion set) generated by the model layer to the business system, complete the "request-response" closed loop.
[0039] In addition to receiving and outputting operations through the above interfaces, the access layer 101 mainly includes independent tools, CI pipeline, TCQ platform, Jingwei platform and other modules.
[0040] Among them, the independent tool is a tool component independent of the main process of the system, which is usually used for quick call of single-point function (such as temporary data processing, tool task execution), provides flexible "tool box" support for business, and meets the non-process, lightweight business needs.
[0041] The CI pipeline is the core carrier of the continuous integration (CI) process, which is responsible for automatically completing code building, testing, deployment and other links, and guarantees code quality and delivery efficiency.
[0042] The TCQ platform generally refers to the technical quality control platform, which focuses on the automatic detection of code quality, technical specifications, test coverage and other dimensions. Through rule checking, static analysis, dynamic testing and other means, technical risks are found in advance, and system stability and maintainability are improved.
[0043] The Jingwei platform is usually used for tracking, analysis and closed-loop management of technical problems (such as fault positioning, performance optimization, technical debt cleaning), which drives the efficient iteration of technical problems from "discovery" to "solution" through data visualization and process automation.
[0044] These modules together constitute the access layer 101, which undertakes the roles of unified access to external needs, pre-guarantee of technical specifications, and flexible supply of tool capabilities. It provides stable, efficient and standardized underlying support for the upper business service layer (such as text preprocessing, rate limiting monitoring, model calling, etc.) and is an indispensable "entry-level" component in the system's "technical foundation".
[0045] (2) Business Service Layer 102: Responsible for business process management. Main contents include: The text preprocessing module performs preprocessing and filtering on the initial text to be detected obtained from model layer 103 to obtain the text to be detected. The preprocessing and filtering operations include: filtering blacklisted content, performing hash calculations, deduplication, formatting (removing leading and trailing spaces), and assembling it into a structured data format that is easy for the system to process. In the rate limiting and frequency control module, the task scheduler manages the distribution and status tracking of detection tasks (such as timeout handling).
[0046] The detection engine in the model call module coordinates the work of modules such as text detection and rule detection. The result processor integrates the results of multiple detection models to generate detection results, which include a set of structured problems (such as syntax errors, semantic contradictions, compliance issues, etc.) and a set of modification suggestions (such as improvement suggestions, replacement solutions, etc.).
[0047] The metric monitoring and data persistence module records detection history and results to the detection log to support source tracing and optimization. The feedback interface module collects user feedback on the detection results (such as "Are the suggestions effective?" and "Have any issues been missed"), thereby optimizing the detection model.
[0048] (3) Model Layer 103: As the core carrier of "detection capability", it achieves multi-level detection of "text to be detected" by combining multiple models. Among them, Model Layer 103 covers the work of multimodal content processing, which is completed by the "text processing workflow", "visual processing workflow" and "file processing workflow" in Dify Flow. Specifically, through the text extraction algorithms in the parsers corresponding to each modality, the text extraction processing of each modality data in the multimodal content is performed to obtain the initial text to be detected; after obtaining the text to be detected, at least two detection models are used to perform basic text detection, semantic and logical detection and compliance detection on the text to be detected in sequence. Moreover, during detection, the detection model adopts a rule detection engine to improve the comprehensiveness and accuracy of detection through a double insurance mechanism.
[0049] Furthermore, since model layer 103 can adapt to the interfaces and capabilities (such as parameter formats and calling methods) of different AI models through model adapters, it can achieve unified calling of multiple models from model vendors. Model vendors may include TMEDeepSeek, HunYuan, Venus, Qwen, etc. The model adaptation layer designs a unified model interface specification, including input format, output format, parameter configuration, etc., and develops dedicated adapters for each large model (DeepSeek / Hunyuan, etc.), which can realize a dynamic discovery and registration mechanism for model capabilities, enabling the text detection and recognition system to flexibly switch between different backend models and optimize resource utilization.
[0050] In addition, since the business service layer 102 collects user feedback on the detection results, the "false alarm analysis workflow" in Dify Flow will provide "experience base" support for the detection model by managing business rules, common error cases and other knowledge through a knowledge base (such as high-frequency problems found in historical detections and industry best practices).
[0051] (4) Independent Service 104: This is the "cornerstone" for ensuring the stable operation of the system, providing underlying support for computing, storage, and monitoring. It mainly covers the following aspects: Computing resources, which provide the necessary GPU / CPU resources for model inference, ensuring good performance and efficiency for large model calls. It can also obtain the text to be detected using plugins, such as Language Tool, PaddleOCR, and Tesseract. Storage system, used to store knowledge base (business rules, cases), detection history (logs, results), and other data to support long-term data accumulation and reuse. Monitoring system, responsible for monitoring system performance (such as response time and resource usage) and operating status (such as service availability), enabling timely detection and handling of faults.
[0052] The text detection and recognition system in this application embodiment can cover diverse scenarios such as pre-publication content inspection, business configuration review, code submission detection, and multimedia content review. It also relies on comprehensive error detection (typo correction, grammatical error recognition, format specification check, semantic error analysis, detection of illegal expressions such as those related to advertising law, and verification of price information accuracy), multi-format content support (single / multiple plain texts, text extraction from image materials, and detection of PDF / CSV / TXT attachments), intelligent modification suggestions (providing modification solutions and explaining the reasons, comparing and recommending multiple solutions), detection report generation (generating detailed reports, classifying and statistically analyzing issues, and supporting report export and sharing), and knowledge base management (knowledge base content query, rule addition and update, and false alarm case record analysis) to achieve multi-dimensional goals of content quality assurance, improved business efficiency, and strengthened compliance management.
[0053] Furthermore, the text detection and recognition system in this embodiment integrates different interfaces to achieve automation. Embedded in the CI / CD process, it automatically performs text checks, bringing text detection to the development stage and significantly reducing process time. Collaborating with the TCQ quality team, it empowers the comprehensive quality inspection process of the business configuration platform. The classification, statistics, export, and sharing functions of the inspection reports allow the quality team to quickly locate problems, track rectification, shorten the quality closed-loop cycle, and enhance cross-departmental collaboration efficiency, truly realizing the value of "automation and collaboration dual-driven" solutions. At the same time, the layered architecture design gives the text detection and recognition system good scalability and maintainability.
[0054] It should be noted that the method provided in this application embodiment can not only be applied to the text detection and recognition system 100 in this application, but can also be developed into a detection tool in the form of a browser plugin. The browser plugin can directly detect the content edited by the user in the browser, provide real-time detection feedback, and is suitable for daily use by operations personnel. It does not require switching tools and provides instant feedback. In addition, it can also be developed as a plugin for mainstream IDEs (such as VSCode and IntelliJ). The IDE plugin can detect text problems in strings in real time when developers are writing code, moving the detection forward to the code writing stage and further reducing the flow of errors.
[0055] In response, this application provides a text detection and recognition method based on the aforementioned text detection and recognition system 100. By using text extraction algorithms corresponding to each modality, text extraction processing is performed on the data of each modality in the multimodal content to obtain the text to be detected. Furthermore, the text to be detected is detected using at least two detection models, thereby improving the detection accuracy and coverage of text detection.
[0056] See Figure 2 The flowchart of a text detection and recognition method provided in this application embodiment is as follows: Figure 2 As shown, the method includes the following steps S201~S204: S201: Receive a text detection request sent by the business system, wherein the text detection request includes multimodal content.
[0057] The business system in this embodiment serves as the external data source and interaction object for the text detection and recognition method, used to initiate detection requests and transmit multimodal content. It should be noted that the business system in this embodiment can be an internal enterprise business system, such as a content management platform (CM system), responsible for generating text (such as advertising copy, product descriptions, and customer service responses) and triggering detection requests; or it can be a third-party business system, such as the backend system of a content review platform, sending requests to the text detection system via an API interface.
[0058] In this step, text detection requests from business systems can be received through various access methods (API, CLI, Web), which lowers the system integration threshold. The text detection request is used to instruct text detection on multimodal content, which can achieve unified detection of multiple formats such as text, images, and PDF.
[0059] In addition, to prevent malicious requests from interfering with the text detection and recognition system, this application embodiment can also verify the text detection requests sent by the business system, such as through signature verification, token verification, IP whitelist and other mechanisms, to ensure that the source of the text detection requests is legitimate.
[0060] S202: Based on the text extraction algorithm corresponding to each modality, text extraction processing is performed on the data of each modality in the multimodal content to obtain the text to be detected.
[0061] Since a copywriting detection request has already been received, this application is able to parse the multimodal content in the copywriting detection request. This multimodal content encompasses text-type data, image-type data, and file-type data. For example, text-type data may include plain text (such as press releases and product descriptions) and structured text (such as tables and comments in code); image-type data may include images (such as posters and text information on product images); and file-type data refers to various PDF / CSV / TXT files.
[0062] In this step, the text extraction algorithm in the corresponding parser is used to extract text from each modality of the multimodal content, thereby obtaining the text to be detected for subsequent detection.
[0063] S203: Detect the text to be detected using at least two detection models and generate detection results, the detection results including a set of structured questions and a set of modification suggestions.
[0064] Since the text to be tested has already been obtained, it can be tested. The specific tests may include performing basic text testing, semantic and logical testing, and compliance testing in sequence. Therefore, for testing at different stages, at least two testing models with different model combinations can be used (for example, a combination of traditional copywriting testing and LLM large model testing). After generating their respective testing results, they are integrated to obtain a set of structured problems (identifying the structural, content, and logical problems in the text) and a set of modification suggestions (providing optimization directions or specific modification solutions for the problems).
[0065] In this step, at least two detection models are used to detect the text to be detected in this embodiment of the application. By utilizing the complementary capabilities of different models, the comprehensiveness and accuracy of the detection are improved. Compared with traditional rule-based detection systems, the detection accuracy and the adoption rate of modification suggestions are improved.
[0066] S204: Send the test results to the business system.
[0067] Since the detection results have already been obtained, they can be transmitted to the business system via an API call, so that the business system can modify and adjust the text in the multimodal content based on the detection results.
[0068] In this step, traditional manual operation requires manual import and export of the content to be tested, which has problems such as cumbersome process and slow information transmission. In this embodiment of the application, the test results can be quickly transmitted to the business system through interface calls, file transfer and other methods, which can replace manual transmission and further reduce time costs.
[0069] As can be seen, this application embodiment receives text detection requests sent by business systems through multiple access methods, achieving flexible integration with different business systems and lowering the integration threshold of the text detection and recognition system. Furthermore, based on the text extraction algorithms corresponding to each modality, text extraction processing is performed on the data of each modality to obtain the text to be detected, achieving unified detection of different formats and avoiding detection omissions caused by format limitations. After obtaining the text to be detected, at least two detection models, such as traditional text detection and LLM large-scale model, can be used to perform multi-stage collaboration including basic text detection, semantic logic detection, and compliance detection, generating detection results containing a set of structured questions and a set of modification suggestions. Compared to traditional rule systems that overly rely on preset rules and specific training corpora, this application embodiment can improve detection accuracy and modification suggestion adoption rate through at least two detection models.
[0070] The following describes one embodiment of the present application, which, compared to the previous embodiment, further explains and optimizes the technical solution. Specifically: S301: Receive a text detection request sent by the business system, wherein the text detection request includes multimodal content.
[0071] S302: Based on the text extraction algorithm corresponding to each modality, text extraction processing is performed on the data of each modality in the multimodal content to obtain the text to be detected.
[0072] As an example, the multimodal content includes text type data, image type data, and file type data. Therefore, step S302 may include: S3021, extracting text information from the text type data using a text extraction algorithm in the text parser; S3022, extracting text information from the image type data using a text extraction algorithm in the image parser; S3023, extracting text information from the file type data by calling the corresponding text extraction algorithm in the file parser based on the file type; S3024, integrating the text information from the text type data, image type data, and file type data to obtain the initial text to be detected; and S3025, preprocessing and filtering the initial text to be detected to obtain the final text to be detected.
[0073] S3021 may include: calling the text extraction algorithm in the text parser to extract text information from text type data (such as plain text files and structured text). The text parser needs to be compatible with different text formats (such as plain text, table / code comments, and other structured text), accurately extract the text type data to be detected in the text, and obtain the text information in the text type data.
[0074] S3022 may include: for image-type data (such as text information in posters or product images), calling an image parser to extract text information. In addition to using text extraction algorithms, the image parser in this embodiment can also utilize multimodal AI models, particularly for intelligent image slicing of long images, solving the technical problem of traditional OCR truncation, thereby recognizing the text embedded in the image and obtaining the text information in the image-type data.
[0075] S3033 can include: for file type data (such as PDF, CSV, TXT, etc.), calling the text extraction algorithm in the corresponding file parser to extract text information based on the file type. For example: for PDF files, the parser needs to parse the document structure (such as chapters, tables, and annotations) and extract text; for CSV files, the parser needs to recognize the table data and extract text; for TXT files, the parser needs to process the plain text format for text extraction. By using the corresponding file parser, it ensures that text information of different file formats is accurately extracted, thus obtaining the text information from the file type data.
[0076] S3034 may include: integrating the three types of text information obtained from the above parsing to form the initial text to be detected. The integration must ensure the integrity and consistency of the three types of information to provide a complete data source for subsequent preprocessing.
[0077] S3035 may include: preprocessing and filtering the initial text to be detected, optimizing the text quality through the following operations: First, filtering is performed, that is, filtering relevant content in the text to be detected through a blacklist, where the blacklist includes common test data, such as keywords such as "test" and "test"; then, noise reduction and standardization are performed, such as removing meaningless characters (such as special symbols and whitespace characters), duplicate text, interference information (such as advertising watermark text), unifying text format (such as capitalization and punctuation), and standardizing text structure (such as paragraph division and word segmentation); finally, it is assembled into a structured data format that is easy for the system to process, thus obtaining the text to be detected.
[0078] In this process, text extraction, integration, and preprocessing are performed on the data of each modality in the multimodal content to achieve full-scene text information extraction from text, images, and files. At the same time, preprocessing improves text quality and lays the foundation for subsequent detection.
[0079] S303: Create and distribute detection tasks based on the copywriting detection request.
[0080] As an example, S303 may include: S3031, creating a detection task based on a text detection request, the detection task being used to instruct the text to be detected to be detected; S3032, determining the task priority corresponding to the detection task based on the type of modal data in the multimodal content and / or the request source of the text detection request; S3033, distributing the detection task to the corresponding detection node based on the type of modal data in the multimodal content and the task priority.
[0081] S3031 may include: after receiving a text detection request, a detection task will be created and a unique task ID will be assigned to each detection task. The task ID serves as a global identifier for the task and is used for subsequent task tracking, status query, processing progress and result association.
[0082] S3032 may include: determining task priority based on the type of modal data (such as text, image, PDF, etc.) in the multimodal content of the copywriting detection request and / or the source of the copywriting detection request (such as the enterprise's internal CM system, third-party review platform, etc.) in combination with business rules or system configuration.
[0083] Furthermore, callers of business systems also support custom priorities and serial / parallel processing. The priority rules are as follows: In terms of content type, key business texts (such as advertisements and product descriptions) have higher priority than regular customer service responses; in terms of request source, core business systems (such as CM systems) have higher priority than third-party platforms; timeliness must also be considered. For example, some sensitive B-end configuration platforms have sequential detection and subsequent business processes (failure to complete detection prevents further steps), requiring high priority due to their timeliness requirements. Some source code detection tasks, due to their business characteristics, require a longer compilation and building time and can be processed in parallel with detection, with relatively lower priority settings.
[0084] Specifically, S3033 can include: requiring the modal data type and task priority in multimodal content, and distributing detection tasks to corresponding detection nodes. For example, if the modal data type is matched as text type data, it is distributed to "text detection nodes" (such as traditional text detection model, Large Language Model (LLM) node); if the modal data type is matched as image type data, it is distributed to "image detection nodes" (such as visual detection model node); if the file type data is matched as file parsing + text detection nodes (such as PDF parsing and text conversion detection nodes). The priority matching is that high-priority tasks are preferentially assigned to idle / resource-sufficient nodes, and low-priority tasks are executed after resources are released.
[0085] During this process, by creating detection tasks, the status and progress of each task can be queried in real time, ensuring traceability and manageability throughout the entire task process. Furthermore, priority settings are used for subsequent task scheduling, ensuring that critical tasks are executed first, improving the rationality of system resource allocation. Additionally, a two-dimensional matching process based on content type and priority ensures that tasks are accurately assigned to nodes with processing capabilities, improving detection efficiency. S304: Based on the detection task, select at least two detection models to detect the text to be detected and generate detection results, which include a set of structured questions and a set of modification suggestions.
[0086] As an example, S304 may include: S3041, selecting at least two detection models to sequentially perform basic text detection, semantic and logical detection, and compliance detection on the text to be detected, to obtain an original set of detection issues; S3042, performing structured sorting based on the error type and severity of the issues in the original set of detection issues, to obtain a structured set of issues; S3043, generating a set of modification suggestions based on the detection models and the structured set of issues.
[0087] It should be noted that the selection of at least two detection models in this application is based on the specific type of the detection task (e.g., plain text detection, image detection, structured data detection) and content features. When no specific detection type is specified, a default detection method will be configured based on the source of the text detection request. For example, plain text detection will be performed by default; if an image external link is found in the multimodal content, image type data detection will be performed separately on that external link.
[0088] As an example, when performing basic text detection, the S3041 can employ a sequential collaboration between traditional text detection and LLM (Large-Scale Model) detection. For instance, traditional detection models (such as rule-based grammar analysis and dictionary-matched spell detection) first perform basic detection on the text to be detected, quickly identifying spelling errors, grammatical errors, punctuation issues, and other problems, generating traditional detection results. These traditional detection results then serve as the pre-context for LLM detection, being passed to the LLM model. The LLM model, combined with this context, performs deeper basic detection on the text to be detected (such as accurate spelling identification based on semantic context and semantic correction suggestions for grammatical errors), generating LLM detection results. The results of traditional and LLM detections are then integrated to form the set of problems identified in the basic text detection stage.
[0089] It should be noted that, in this embodiment, when selecting a large language model to detect the text, a sliding window mechanism is used to divide the text into multiple context-related text segments. The large language model performs long-text understanding on each text segment and combines this with the contextual information to perform error detection, thus obtaining the LLM detection result. Since traditional detection methods are mostly based on single-sentence analysis and ignore contextual relationships, this embodiment uses an LLM large model to achieve context-aware error detection capabilities, which can improve detection efficiency and accuracy.
[0090] As an example, after performing basic text detection, S3041 selects a detection model to perform semantic and logical detection on the text to be detected, mainly detecting semantic coherence, business logic, price consistency, etc. Similarly, it can also use the LLM large model and other models in combination for detection to analyze whether the text semantics are coherent and identify unclear or contradictory content; based on the business knowledge base (such as industry standards and enterprise processes), it checks whether the copy description conforms to business logic, and checks whether the price information in the text to be detected is consistent and conforms to business rules, forming a set of problems identified in the semantic and logical detection stage.
[0091] As an example, after the first two layers of detection, S3041 selects a detection model to perform compliance checks on the text to be detected, mainly checking for compliance with advertising laws, sensitive content, brand guidelines, etc. Similarly, it can also use the LLM large model in combination with other models to detect and identify expressions that violate advertising laws (such as absolute terms such as "best" and "superior"); identify content that may cause controversy or inappropriateness; and check whether brand names, trademarks, etc. are used in accordance with regulations, forming a set of issues identified in the compliance detection stage.
[0092] Specifically, S3042 may include: after obtaining the original set of detected issues, performing a structured sorting based on the error type and severity of the issues in the original set. The error type and severity are preset by the system; for example, severity levels include suggestions, warnings, and errors; while error types include spelling errors, grammatical errors, and logical errors. Since different error types and severity levels have different weights and priorities, a structured sorting can be performed based on these factors. However, different error types and severity levels may be adjusted according to different business characteristics; different businesses may assign different weights to "severity" (e.g., in advertising copy, "advertising law violation" might be directly classified as "error," while in technical documents, "spelling error" might only be classified as "suggestion").
[0093] In addition, this application embodiment designs a dedicated prompt template for each type of error (such as spelling, grammar, and logic), clearly defining the detection logic and output format. Furthermore, based on business characteristics, the prompt is dynamically generated and continuously optimized. For example, the advertising copy prompt needs to include "advertising law compliance detection" logic, and the technical document prompt needs to include "technical terminology accuracy detection" logic. The prompt is also iterated based on the detection results: if a certain type of error is missed / falsely detected, details such as the "detection logic description" and "output format constraints" in the prompt are adjusted to improve detection accuracy and continuously optimize the prompt effect.
[0094] Specifically, S3043 can include generating a set of modification suggestions based on the detection model and the structured question set, considering three dimensions: readability analysis, concise expression, and style consistency. This set of modification suggestions provides specific suggestions, explanations of the reasons and basis for the modifications, and comparisons and recommendations of multiple solutions. It should be noted that readability analysis is used to evaluate and optimize text fluency, concise expression is used to identify and optimize redundant expressions, and style consistency is used to check and unify text style. In this process, readability analysis ensures fluency, concise expression improves efficiency, and style consistency enhances professionalism, thereby optimizing the quality of the modification suggestion text and ultimately generating a targeted and logically clear set of modification suggestions.
[0095] It should be noted that, in this embodiment of the application, the test results can be stored in a database, which provides a centralized storage and systematic management platform for the test results. This allows business personnel to perform batch operations (such as batch importing / exporting test data) and quickly query historical results, avoiding data dispersion and loss.
[0096] This process supports model combination and invocation, leveraging the strengths of different large AI models in detecting various types of errors to build a multi-model collaborative detection mechanism. This allows for the integration of detection results from multiple models. Furthermore, it enables load balancing and failover for model invocation, thereby improving detection efficiency.
[0097] S305: Send the test results to the business system.
[0098] S306: Obtain the detection result feedback sent by the business system.
[0099] Once the business system receives the detection results, it can provide feedback on the results, such as choosing to "accept" or "reject" modification suggestions, and can supplement the feedback reasons (such as the reason for rejection, the optimization direction after acceptance). In order to optimize the overall copywriting detection and recognition method, it is necessary to collect false positive and false negative cases. Therefore, after the business system provides feedback on the detection results and generates the detection result feedback, it needs to be sent to the copywriting detection and recognition system.
[0100] S307: Perform data analysis on the test results feedback and optimize the error correction knowledge base based on the data analysis results.
[0101] In this process, this application embodiment requires data analysis of the detection result feedback. Based on the user feedback scenario of "rejecting modification suggestions," misjudged fault modes (such as being detected as A but actually being B) are added to the knowledge base as "error modes." Combined with business rule vulnerabilities discovered in the data analysis results (such as the detection logic for a certain type of fault mode not matching actual business practices), the business rules and terminology relationships in the knowledge base are updated. Furthermore, an error correction knowledge base update mechanism can be established to automatically or manually update the knowledge base content periodically based on the detection result feedback and data analysis results, ensuring its continuous adaptation to business scenarios.
[0102] S308: Adjust the model weights of the detection model based on the feedback of detection results and the evaluation of historical detection effects.
[0103] As an example, S308 may include: evaluating the model capability of the detection model based on the detection results feedback and historical detection performance to obtain a model score, wherein the historical detection performance includes core indicators such as accuracy, recall, and false negative rate of the detection model output; and dynamically adjusting the model weights corresponding to the detection model based on the model score.
[0104] The specific process of evaluating the capabilities of detection models may include: assigning scoring weights to each detection model from multiple dimensions based on the requirements of the detection task, and calculating the model score of the detection model by combining the feedback of detection results and historical performance through methods such as weighted averaging and analytic hierarchy process.
[0105] In this process, by dynamically adjusting the weights, the detection model's multi-dimensional capabilities are enhanced in a targeted manner, significantly improving detection accuracy. Furthermore, based on feedback from detection results and historical detection performance, the detection model can quickly adapt to changes in business scenarios, ensuring that analysis and decision-making always align with business needs.
[0106] To address this, a complete feedback loop mechanism can be established to continuously collect false alarm cases and optimize the system. Furthermore, combined with the TCQ quality team's control process, a full-process quality assurance system can be formed from development to release.
[0107] In summary, this application's embodiments, upon receiving a multimodal detection request, parse the content and extract text from unstructured documents such as images and PDFs, transforming multimodal information into a detectable text stream. It employs at least two detection models for text detection, integrating the complementary advantages of different models to accurately identify complex semantic errors, business logic errors, and violations that are difficult to detect using traditional methods. It generates detection results containing a set of structured questions and a set of modification suggestions, providing a high-quality data foundation for model optimization. Simultaneously, it defines a clear request-response interface, service-orientedizing the detection capabilities and directly integrating them into other business systems. This achieves efficient parsing of multimodal content, accurate identification of complex errors, and deep adaptation to business processes, providing a systematic solution for text detection scenarios.
[0108] It should be noted that the text detection and recognition system in this application is mainly applied to the following scenarios: (1) Pre-release content check: Conduct a comprehensive check before releasing operational content, product introductions, activity descriptions and other texts to avoid launching incorrect content.
[0109] (2) Business configuration review: Check business configuration data in JSON and other formats to ensure that the configuration is accurate.
[0110] (3) Code submission detection: Integrated into the CI / CD process, automatically detects textual issues in the source code during the code submission stage.
[0111] (4) Multimedia content review: Extract and detect text in multimedia content such as pictures and PDFs.
[0112] (5) Quality control process: Collaborate with the TCQ quality team to achieve a comprehensive quality inspection process covering the business configuration platform.
[0113] The following describes how the text detection and recognition method of this application embodiment detects and recognizes different types of content and generates detection results in the above application scenarios.
[0114] Among them, when checking the application scenarios of the copy before content publication, the specific detection results for the corresponding text types are as follows: Figure 4 As shown, basic text detection can identify errors such as misspelled words, missing words, punctuation, and terminology. Semantic and logical detection can uncover grammatical errors. Compliance detection identifies issues related to pricing, song titles, and other compliance matters. Furthermore, the system generates corresponding modification suggestions for each detected problem. For example, for pricing errors, the suggestion is to set the Green Diamond VIP price within the range of 5-18 yuan per month. In other words, the system generates corresponding modification suggestions for each type of error, providing users with a clear and concise approach to improvement.
[0115] In application scenarios involving business configuration review and code submission detection, the text type corresponding to the code format will be detected. The specific detection results are as follows: Figure 5 As shown. Basic text detection can detect English spelling errors; semantic and logical detection can identify mismatches between singer IDs and songs, as well as mismatches in business code; compliance detection can identify abnormal link formats. Furthermore, the system generates corresponding modification suggestions for the detected issues. For example, for an abnormal link format error, the modification suggestion is to change it to https: / / y.qq.com.
[0116] In multimedia content moderation applications, image and file type data (PDF format content) will be detected accordingly. The specific detection results are as follows: Figure 6 As shown. Basic text detection can identify English spelling errors and misspelled / omitted words; semantic and logical detection can identify date and numbering errors; and compliance detection can identify incorrect names / movie titles. Furthermore, the system generates corresponding modification suggestions for the detected issues. For example, for non-existent date formats, such as June not having a 31st, it should be changed to another reasonable date.
[0117] The text detection and recognition scheme proposed in this embodiment demonstrates significant value in multiple scenarios: before content publication, it performs basic text, semantic logic, and compliance checks on operational copy and product introductions, and generates structured questions and modification suggestions, providing quality assurance for content launch, business configuration, code submission, and multimedia review. See Figure 7This application provides a text detection and recognition device 700, which includes: The receiving unit 701 is used to receive a text detection request sent by the business system, wherein the text detection request includes multimodal content; The parsing unit 702 is used to perform text extraction processing on the data of each modality in the multimodal content based on the text extraction algorithm corresponding to each modality, so as to obtain the text to be detected; The detection unit 703 is used to detect the text to be detected using at least two detection models and generate detection results, the detection results including a set of structured questions and a set of modification suggestions; The sending unit 704 is used to send the detection results to the business system.
[0118] In a preferred embodiment, the multimodal content includes text type data, image type data, and file type data, and the detection unit 703 is specifically used for: Extract text information from text-type data using text extraction algorithms in a text parser; Extract text information from image data using a text extraction algorithm in an image parser; Based on the file type, the text extraction algorithm in the corresponding file parser is invoked to extract text information from the file type data; By integrating text information from text-type data, text information from image-type data, and text information from file-type data, the initial text to be detected is obtained; The initial text to be detected is preprocessed and filtered to obtain the text to be detected.
[0119] In a preferred embodiment, the document detection and recognition device 700 further includes: a task distribution unit, having functions for: A detection task is created based on a text detection request. The detection task is used to instruct the text to be detected to be detected. Determine the task priority corresponding to the detection task based on the type of modal data in the multimodal content and / or the source of the copy detection request; Based on the type of modal data in the multimodal content and the task priority, the detection task is distributed to the corresponding detection node.
[0120] In a preferred embodiment, the detection unit 703 is specifically used for: Select at least two detection models to sequentially perform basic text detection, semantic and logical detection, and compliance detection on the text to be detected, and obtain the original set of detection questions; The original set of detection issues is sorted in a structured manner according to the error type and severity of the issues to obtain a structured set of issues; Based on the detection model and the set of structured questions, a set of modification suggestions is generated.
[0121] In a preferred embodiment, the detection unit 703 is also used for: When selecting a large language model to detect text, a sliding window mechanism is used to divide the text into multiple context-related text segments. The large language model performs long-text understanding on each text segment and performs error detection in combination with the context to obtain the detection results of the large model.
[0122] In a preferred embodiment, the document detection and recognition device 700 further includes an optimization unit specifically used for: Obtain the detection results feedback sent by the business system; Data analysis is performed on the detection results, and the error correction knowledge base is optimized based on the data analysis results. The error correction knowledge base is built based on a business terminology database and business rules to enhance the business capabilities of the detection model.
[0123] In a preferred embodiment, the document detection and recognition device 700 further includes: an adjustment unit specifically used for: The model's capabilities are evaluated based on the feedback of detection results and historical detection performance to obtain a model score; The model weights corresponding to the detection model are dynamically adjusted based on the model score.
[0124] It should be noted that the specific implementation method and the effects achieved by the text detection and recognition device 700 can be found in the above. Figure 2 or Figure 3 The relevant descriptions in the provided methods will not be repeated here.
[0125] This application also provides an electronic device, see [link to document]. Figure 8 The present application provides a structural diagram of an electronic device 80, as shown in the embodiment. Figure 8 As shown, it may include a processor 81 and a memory 82.
[0126] The processor 81 may include one or more processing cores, such as a core processor or a core processor. The processor 81 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 81 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 81 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 81 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0127] The memory 82 may include one or more computer-readable storage media, which may be non-transitory. The memory 82 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory 82 is used to store at least the following computer program 821, which, after being loaded and executed by the processor 81, is capable of implementing the relevant steps in the document detection and recognition generation method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 82 may also include an operating system 822 and data 823, etc., and the storage method may be temporary storage or permanent storage. The operating system 822 may include Windows, Unix, Linux, etc.
[0128] In some embodiments, the electronic device 80 further includes a display screen 83, an input / output interface 84, a communication interface 85, a sensor 86, a power supply 87, and a communication bus 88.
[0129] certainly, Figure 8 The structure of the electronic device shown does not constitute a limitation on the electronic device in the embodiments of this application. In practical applications, the electronic device may include more than [other components]. Figure 8 More or fewer components as shown, or combinations of certain components.
[0130] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the text detection and recognition steps performed by the electronic device of any of the above embodiments.
[0131] In another exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the text detection and recognition steps performed by the electronic device of any of the above embodiments.
[0132] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the methods of the above embodiments can be implemented by means of software plus a general-purpose hardware platform. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as a read-only memory (ROM) / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, a server, or a network communication device such as a router) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0133] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. Modules described as separate components may or may not be physically separate. Components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the objectives of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0134] The above description is merely an exemplary implementation of this application and is not intended to limit the scope of protection of this application.
Claims
1. A text detection and recognition method, characterized in that, The method includes: Receive a text detection request sent by the business system, the text detection request including multimodal content; Based on the text extraction algorithms corresponding to each modality, text extraction processing is performed on the data of each modality in the multimodal content to obtain the text to be detected; The text to be detected is detected using at least two detection models to generate detection results, which include a set of structured questions and a set of modification suggestions. The detection results are sent to the business system.
2. The method according to claim 1, characterized in that, The multimodal content includes text data, image data, and file data. The text extraction algorithm corresponding to each modality performs text extraction processing on each modality data in the multimodal content to obtain the text to be detected, including: The text information in the text type data is extracted using a text extraction algorithm in a text parser; The text information in the image type data is extracted using the text extraction algorithm in the image parser; Based on the file type, the text extraction algorithm in the corresponding file parser is invoked to extract the text information from the file type data; By integrating the text information from the text type data, the text information from the image type data, and the text information from the file type data, an initial text to be detected is obtained; The initial text to be detected is preprocessed and filtered to obtain the text to be detected.
3. The method according to claim 2, characterized in that, Before performing text extraction processing on the modal data in the multimodal content, the text extraction algorithm based on each modality respectively includes: A detection task is created based on the text detection request, and the detection task is used to instruct the text to be detected to be detected. The task priority corresponding to the detection task is determined based on the type of modal data in the multimodal content and / or the request source of the text detection request; Based on the type of modal data in the multimodal content and the task priority, the detection task is distributed to the corresponding detection node.
4. The method according to claim 1, characterized in that, The step of detecting the text to be detected using at least two detection models and generating detection results includes: Select at least two detection models to sequentially perform basic text detection, semantic and logical detection, and compliance detection on the text to be detected, and obtain the original set of detection questions; The original set of detected problems is sorted in a structured manner according to the error type and problem severity to obtain the structured problem set. Based on the detection model and the set of structured questions, the set of modification suggestions is generated.
5. The method according to claim 4, characterized in that, The large language model is selected to detect the text to be detected, including: The text to be detected is divided into multiple context-related text segments using a sliding window mechanism; The large language model is used to perform long-text understanding on each text segment, and error detection is performed in combination with the context to obtain the detection results of the large model.
6. The method according to claim 1, characterized in that, After sending the detection results to the business system, the process further includes: Obtain the detection result feedback sent by the business system; The detection results are analyzed, and the error correction knowledge base is optimized based on the analysis results. The error correction knowledge base is built on a business terminology database and business rules to enhance the business capabilities of the detection model.
7. The method according to claim 6, characterized in that, The method further includes: The detection model's capabilities are evaluated based on the feedback from the detection results and historical detection performance, resulting in a model score. The model weights corresponding to the detection model are dynamically adjusted based on the model score.
8. A document detection and recognition system, characterized in that, The system includes: an access layer, a business service layer, and a model layer; The access layer is used to receive text detection requests sent by the business system, and the text detection requests include multimodal content; The model layer is used to perform text extraction processing on the data of each modality in the multimodal content based on the text extraction algorithm corresponding to each modality, so as to obtain the text to be detected; and to detect the text to be detected by at least two detection models. The business service layer is used to integrate the results of the detection model to generate detection results, which include a set of structured questions and a set of modification suggestions. The access layer is also used to send the detection results to the business system.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the text detection and recognition method as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the text detection and recognition method as described in any one of claims 1 to 7.