Handwriting experiment sample intelligent extraction system

The intelligent handwriting sample extraction system automatically collects and generates handwriting content that meets the identification requirements and verifies its authenticity. This solves the problems of low quality handwriting sample extraction and low identification efficiency in existing technologies, and achieves efficient and accurate handwriting sample management.

CN121582946APending Publication Date: 2026-02-27HUBEI POLICE ACAD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511557656.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

In existing technologies, the extraction of handwriting test samples is often poorly performed by non-professionals, resulting in low quality that fails to meet the requirements for identification. There is also a lack of means to confirm authenticity, and data entry is time-consuming and inaccurate, which increases the complexity and efficiency of case identification.

Method used

A handwriting experiment sample intelligent extraction system is provided, including an information acquisition module, a handwriting generation module, and a sample verification module. The system automatically collects data through integrated hardware devices, generates handwriting content that meets the identification requirements using an NLP model, and verifies its authenticity and integrity.

Benefits of technology

It enables standardized, efficient, and accurate extraction of handwriting samples, reduces manual intervention, improves data entry speed and accuracy, and ensures the authenticity of samples and the quality of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121582946A_ABST
    Figure CN121582946A_ABST
Patent Text Reader

Abstract

The invention discloses a handwriting experiment sample intelligent extraction system, and relates to the field of material evidence investigation and identification, and the system comprises an information collection module which is used for collecting related data and generating structured data; the related data comprises personnel information, case information and environment data; the handwriting generation module is used for processing the structured data and generating handwriting content meeting the identification requirement; the handwriting content comprises a handwriting text and a voice file; and the sample confirmation module is used for carrying out integrity confirmation on the handwriting content and carrying out authenticity verification on the extraction person and the extracted person to generate a confirmation result. The handwriting experiment sample extracted by the method can meet the identification requirement, and the case identification quality and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of physical evidence examination and identification, and in particular to an intelligent extraction system for handwriting experimental samples. Background Technology

[0002] With the rapid development of science and technology, information technology has been widely used in various industries. However, in practice, handwriting sample extraction faces many challenges, such as non-standard extraction content and operation, resulting in low sample extraction quality and failure to meet basic identification requirements. At present, the focus is still on manual extraction by professional technicians and continuous improvement of manual extraction techniques. No intelligent extraction system or equipment for handwriting samples has yet been found.

[0003] The main manifestations are: Firstly, the extraction by non-professional technicians resulted in fewer identical characters in the experimental samples, failing to meet the basic requirements for identification; secondly, the extraction by copying methods resulted in samples that could not properly reflect writing habits, failing to meet the basic identification conditions, thus requiring repeated extraction. Secondly, traditional handwriting sample collection often lacks effective means of authenticity verification and traceability mechanisms, which not only easily leads to questions about the legality of evidence but also increases the difficulty of subsequent verification and investigation. At the same time, there are also pain points and difficulties in handwriting experimental sample extraction, such as time-consuming information entry, low accuracy, samples failing to meet identification requirements, and weak data traceability capabilities. This not only increases the complexity of handwriting examination work but also affects the quality and efficiency of case identification. Summary of the Invention

[0004] The purpose of this application is to provide an intelligent handwriting sample extraction system to solve the problem that handwriting samples extracted by non-professionals cannot meet the identification requirements and the quality and efficiency of case identification are low.

[0005] To achieve the above objectives, this application provides the following solution: Firstly, this application provides an intelligent extraction system for handwriting test samples, comprising: The information collection module is used to collect relevant data and generate structured data; the relevant data includes personnel information, case information, and environmental data. The handwriting generation module is used to process the structured data and generate handwriting content that meets the identification requirements; the handwriting content includes handwriting text and audio files; The sample verification module is used to verify the integrity of the handwriting content and the authenticity of the extractor and the extracted person, and generate a verification result.

[0006] According to the specific embodiments provided in this application, this application has the following technical effects: This application constructs an intelligent handwriting sample extraction system. Through an information acquisition module, a handwriting generation module, and a sample confirmation module, it generates handwriting content that meets the identification requirements. This handwriting content includes handwriting text and audio files, realizing standardized, efficient, and accurate extraction and management of handwriting samples. Through automated information entry and sample text generation functions, it significantly reduces manual intervention and improves the speed and accuracy of data entry. Finally, the sample confirmation module confirms the integrity of the handwriting content and verifies the authenticity of the extracted person, thereby ensuring the authenticity of the handwriting samples and improving the quality and efficiency of case identification. Attached Figure Description

[0007] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 This is a schematic diagram of the intelligent extraction system for handwriting experimental samples provided in an embodiment of this application; Figure 2 This is a schematic diagram of the information acquisition module processing flow provided in one embodiment of this application; Figure 3 This is a schematic diagram of the handwriting generation module processing flow provided in an embodiment of this application; Figure 4 This is a schematic diagram of the sample verification module processing flow provided in an embodiment of this application; Figure 5 This is a schematic diagram of the data management module processing flow provided in an embodiment of this application. Detailed Implementation

[0009] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0010] To make the objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0011] like Figure 1 As shown in the figure, this application provides an intelligent extraction system for handwriting test samples, including: The information collection module is used to collect relevant data and generate structured data; the relevant data includes personnel information, case information, and environmental data.

[0012] The handwriting generation module is used to process the structured data and generate handwriting content that meets the identification requirements; the handwriting content includes handwriting text and audio files.

[0013] The sample verification module is used to verify the integrity of the handwriting content and the authenticity of the extracted person, and generate a verification result.

[0014] In one exemplary embodiment, such as Figure 2 As shown, the information collection module specifically includes: A personnel information collection unit is used to collect personnel information using integrated hardware devices. These integrated hardware devices include an ID card reader, a fingerprint sensor, a signature pad, and a GPS locator. The ID card reader is used to automatically extract the ID card information of the person being collected using OCR technology. The fingerprint sensor is a capacitive fingerprint sensor used to collect the biometric features of the person being collected and generate an encrypted fingerprint template. The signature pad is used to capture the signature trajectory of the person being collected and store it as a vector graphic. The GPS locator is used to automatically obtain the coordinates of the collection location. The personnel information includes the ID card information, the encrypted fingerprint template, the vector graphic, and the coordinates of the collection location.

[0015] The case information collection unit is used to generate a unique extraction number based on the timestamp and random number according to the input fields, and to collect the case information according to the extraction number; the input fields include case number, client, case name, case type, extraction location, detailed address and submission time.

[0016] An environmental data acquisition unit is used to automatically record environmental data using equipment sensors; the environmental data includes timestamps, equipment IDs, operator IDs, GPS coordinates, and network status.

[0017] The structured data generation unit is used to generate structured data from the personnel information, the case information, and the environmental data.

[0018] In practical applications, the information acquisition module: Location: Located in the system user interaction layer.

[0019] Function: Responsible for collecting personnel information, case information, and environmental data to provide basic input for subsequent handwriting generation. It achieves efficient and accurate data collection through automated tools (such as ID card readers, fingerprint sensors, and GPS locators).

[0020] Connection Relationship: As the data input source, its output (structured data) is transmitted to the handwriting generation module via the system's internal bus or API interface, forming a unidirectional data flow. This module is indirectly associated with the sample verification module and shares the acquisition results through the data management module.

[0021] Data collection method: Personnel information collection: Automated collection is performed using integrated hardware devices (such as ID card readers, fingerprint sensors, and signature pads). Specific methods include: ID card recognition: Automatically extracts fields such as name, ID number, registered address, and gender using OCR technology, reducing manual input errors.

[0022] Fingerprint registration: A capacitive fingerprint sensor is used to collect biometric features and generate an encrypted fingerprint template (represented as binary data).

[0023] Electronic signature: Capture the signature trace via touch screen or signature pad and store it as a vector graphic (SVG format).

[0024] GPS positioning: Automatically acquires the coordinates (latitude and longitude) of the data collection location, and displays them in WGS84 format.

[0025] Case information collection: Manual input via user interface (UI) forms, combined with automatic generation functionality: Input fields: Case number, client, case name, case type, collection location, detailed address, submission time, etc.

[0026] Automatic generation: Click the "Get Extraction Number" button, and the system will generate a unique extraction number based on the timestamp and a random number.

[0027] Environmental data acquisition: Automatically record timestamps, device IDs, network status, etc. using device sensors (such as GPS and network modules), and represent them in JSON format.

[0028] Information collected: Personnel information: including information of the person extracting the information (name, identity category, electronic signature) and information of the person writing the information (name, ID number, registered address, gender, contact information, fingerprint template, electronic signature).

[0029] Case information includes case number, client, case name, case type, collection location, detailed address, submission time, and collection number.

[0030] Environmental data includes timestamps (ISO 8601 format), device ID, GPS coordinates, and operator ID.

[0031] Information Representation and Output: All information is represented as structured data (JSON format). The output is to transmit the structured data to the handwriting generation module via the internal API, and at the same time back it up to the temporary cache of the data management module.

[0032] In one exemplary embodiment, such as Figure 3 As shown, the handwriting generation module specifically includes: The handwriting text generation unit is used to preprocess the structured data by calling an NLP model based on the custom keyword content, and generate handwriting text that conforms to the context.

[0033] The speech file generation unit is used to convert the handwritten text into a speech file.

[0034] The handwriting content generation unit is used to generate handwriting content based on the handwriting text and the audio file.

[0035] In practical applications, the handwriting generation module: Location: Located in the system's business logic layer.

[0036] Functionality: Based on input data from the information acquisition module, this system generates handwriting text and audio files that meet authentication requirements through Natural Language Processing (NLP) and speech synthesis technologies. It supports both local and cloud-based generation modes to ensure content diversity and real-time performance.

[0037] Connection relationship: Receives output data from the information acquisition module, processes it, and transmits the generated handwriting content (text + voice) to the sample confirmation module. It interacts bidirectionally with the data management module, calling up lexicon / dictionary data and storing the generated results.

[0038] In an exemplary embodiment, the handwriting text generation unit specifically includes: The generation mode selection subunit is used to select the generation mode based on the custom keyword content; the generation modes include local mode and cloud mode.

[0039] The local mode handwriting text generation subunit is used to invoke a local NLP model and use a pre-trained word-filling and sentence-building expansion model to generate handwriting text that conforms to the context when the generation mode is local mode; the local NLP model is a lightweight model and the lightweight model is the ONNX Runtime framework.

[0040] The cloud-based handwriting text generation subunit is used to call a Transformer-based cloud NLP model to generate handwriting text that conforms to the context when the generation mode is cloud-based; the cloud NLP model includes a Chinese model and an English model.

[0041] In practical applications, the processing logic is as follows: based on the input keywords, handwritten text and audio files are generated via local or cloud-based methods. The processing flow includes: 1. Keyword Input: Enter custom keywords. Keywords support the " / " separator to separate phrases (e.g., "contract / signature" represents two independent phrases). 2. Text Generation: Based on keywords, an NLP model is invoked to generate handwritten text that fits the context. The generation process supports word count control (user inputs the number of words to be generated).

[0042] 3. Speech Synthesis: Converts generated text into speech files (WAV format), supporting multiple languages ​​(Chinese / English) and playback control (speed adjustment, loop playback).

[0043] Generation method: Local generation: A lightweight model running on the Android platform (such as the ONNX Runtime framework) using a pre-trained word-filling and sentence-building model. Input consists of keywords (strings) and the number of words to be generated (integers). Output consists of handwritten text (strings) and audio files (WAV format). It is characterized by offline operation, and response speed depends on device performance (typical response time <5 seconds).

[0044] Cloud-based generation: Utilizes high-performance NLP models (such as the T5_Mask_Completion Chinese model and the genius-large English model) via cloud API calls. Input consists of keywords (strings), language mode (enumeration: Chinese / English), and the number of characters to be generated (integer). Output consists of handwritten text (strings) and audio files (WAV format). Features include richer text (trained on a dataset of 80 million clean paragraphs) and faster response time (typical response time <2 seconds).

[0045] In practical applications, the text generation algorithm (based on the Transformer NLP model) is shown below.

[0046] Algorithm selection: Cloud-based text generation employs Transformer-based NLP models, specifically T5_Mask_Completion (Chinese) and genius-large (English). This model generates context-appropriate handwriting text through Masked Language Modeling (MLM).

[0047] Model structure: Overall architecture: standard Transformer encoder-decoder structure, including multi-head self-attention mechanism and feed-forward neural network.

[0048] Key components: Encoder: Converts the input keyword sequence into a context vector. Formula: EncoderOutput = LayerNorm(X + MultiHeadAttention(X,X,X)) where (X) is the input embedding vector, and (MultiHeadAttention) calculates the attention weights.

[0049] Decoder: Generates the target text based on the encoder output. Formula: DecoderOutput = LayerNorm(Y + MultiHeadAttention(Y, EncoderOutput, EncoderOutput)) where (Y) is the decoder input.

[0050] Masking mechanism: During the training phase, 15% of the tokens in the input sequence are randomly masked, and the model predicts the masked tokens.

[0051] Model parameters: The T5_Mask_Completion model has approximately 220 million parameters and supports a maximum input length of 512 tokens; the genius-large model has approximately 700 million parameters and supports multilingual generation.

[0052] Input / output: Input: Keyword string (e.g., "contract / signature"), number of characters to generate (integer), language mode (enumeration: Chinese / English).

[0053] Output: Handwriting text string (length controlled by the number of characters generated).

[0054] Training process: Dataset preparation: Pre-training was performed using a dataset of 80 million clean Chinese paragraphs (generated in the cloud) and a local lightweight dataset (100,000 paragraphs).

[0055] Pre-training: Perform the MLM task on a large-scale text corpus, with the optimization objective being to minimize the cross-entropy loss of the mask tokens.

[0056] Fine-tuning: The model is fine-tuned using a dedicated dataset of handwriting samples (including keyword-text pairs) to optimize the fluency and relevance of the generated text.

[0057] Evaluation: The quality of the generated product is evaluated using the BLEU-4 and ROUGE metrics, with BLEU-4 required to be greater than 0.8.

[0058] Training environment: cloud-based GPU cluster (such as NVIDIA V100), training time approximately 48 hours.

[0059] Application process: Inference phase: The user inputs keywords and word count, and the system calls the model API.

[0060] Text generation: The model processes keywords through an encoder, and the decoder generates text through autoregression (maximum generated length = number of generated characters).

[0061] Post-processing: Filter the generated text (e.g., remove sensitive words) to ensure compliance with specifications.

[0062] Output: Returns the handwriting text to the handwriting generation module.

[0063] The handwriting generation module combines cloud data processing with local generation, enabling flexible switching between different application scenarios and ensuring the diversity and quality of handwritten text generation. Simultaneously, by introducing electronic signature and fingerprint recognition technologies, combined with sample file uploads and timestamp recording, the authenticity, legality, and traceability of handwriting samples are ensured, meeting the high standards required for authentication.

[0064] In one exemplary embodiment, the voice file generation unit specifically includes: The post-processing subunit is used to perform post-processing on the handwriting text and determine the processed handwriting text.

[0065] The speech file generation subunit is used to perform speech synthesis using TTS technology to generate speech files.

[0066] In one exemplary embodiment, such as Figure 4 As shown, the sample confirmation module specifically includes: The sample uploading unit is used to acquire handwriting sample images collected on-site.

[0067] The keyword matching degree checking unit is used to perform keyword matching degree checking between the handwriting sample image and the handwriting content, determine whether the keyword matching degree is greater than the matching degree threshold, and determine the first judgment result.

[0068] The integrity verification unit is used to verify the integrity of the handwriting content based on the digital fingerprint of the handwriting sample image.

[0069] The biometric verification unit is used to perform biometric verification on the extracted person based on the personnel information in the information collection module, determine whether the biometric verification is successful, and confirm the second judgment result.

[0070] The confirmation result generation unit is used to determine that the authenticity of the extracted person has been verified, record the current timestamp, generate a UTC timestamp, bind the handwriting sample image with the handwriting content, and generate a confirmation result.

[0071] The return unit is used to return to the information collection module, re-collect relevant data, and generate structured data.

[0072] In practical applications, the sample verification module: Location: Located in the system verification layer, it operates after the handwriting generation module to confirm the accuracy of each sample data.

[0073] Function: Verify the authenticity and completeness of the handwriting content output by the handwriting generation module, and ensure the legality and validity of the sample through biometric technology (electronic signature, fingerprint entry).

[0074] Connection relationship: Receives the output from the handwriting generation module, processes it, and transmits the confirmation result (including timestamp, signature, etc.) to the data management module for storage. This forms a closed-loop feedback with the information acquisition module; if confirmation fails, it can trigger a re-acquisition.

[0075] Function: By uploading on-site photos, the number of materials, and collecting the writer's biometric information (confirming electronic signatures and fingerprints), the authenticity and integrity of the handwriting content output by the handwriting generation module are verified to ensure the legality and validity of the samples.

[0076] In one exemplary embodiment, the biometric verification unit specifically includes: The integrity verification subunit is used to calculate the digital fingerprint of the handwriting sample image using a hash algorithm, and to verify the integrity of the handwriting content based on the digital fingerprint.

[0077] The input subunit is used to input the fingerprint information of the person being extracted.

[0078] The similarity calculation subunit is used to calculate the similarity between the fingerprint information and the fingerprint template in the information acquisition module using the Hamming distance algorithm.

[0079] The authenticity verification subunit is used to determine that the authenticity verification of the extracted person has passed when the similarity is greater than a set threshold.

[0080] In practical applications, the sample verification module's processing logic is as follows: This module receives the output (handwriting text and audio files) from the handwriting generation module and verifies the authenticity and integrity of the samples. The processing flow includes: Sample Upload: Users upload handwriting sample images (JPG / PNG format) collected on-site by taking photos with their cameras.

[0081] Information verification: Obtain the number of handwriting samples collected on-site and generate a number of style confirmation copies, which are then provided to the writer for confirmation.

[0082] Biometric verification: The writer is required to make an electronic signature and enter their fingerprint using a signature pad and fingerprint sensor to verify their identity.

[0083] Timestamp recording: Automatically obtains the current time (ISO 8601 format) and generates a unique timestamp.

[0084] Confirmation method: Electronic signature: Capture the signature trace using a touchscreen to generate an encrypted signature file (format: PKCS#7).

[0085] Fingerprint verification: Compare the entered fingerprint with the fingerprint template in the information collection module, and use the Hamming distance algorithm to calculate the similarity (threshold >90% is considered as passing).

[0086] Sample integrity check: The digital fingerprint of the sample file is calculated using a hash algorithm (SHA-256) to ensure that it has not been tampered with.

[0087] In one exemplary embodiment, such as Figure 5 As shown, this application also includes: a data management module for managing the data storage, retrieval, backup, and security management of the handwriting experimental sample intelligent extraction system.

[0088] In practical applications, the data management module includes: Location: Located in the system data layer, it is the system's storage and management center.

[0089] Functions: Responsible for the storage, retrieval, backup, and security management of all system data, including information collection, content generation, and confirmation records. Supports resource export and access control to ensure data traceability.

[0090] Connection relationships: It is associated with all three modules mentioned above: receiving raw data from the information collection module, the generated results from the handwriting generation module, and the verification records from the sample confirmation module, and providing a unified data interface for all modules to call. Data is stored in an SQLite database on the device terminal via local service communication.

[0091] In one exemplary embodiment, the data management module specifically includes: The storage unit is used to store different types of data from the intelligent extraction system of handwriting experimental samples and to synchronize different types of data to the cloud server.

[0092] The retrieval unit supports multi-condition retrieval and returns case details; the multi-condition retrieval includes case number retrieval and time range retrieval; the case details include all related data related to the case.

[0093] The resource management unit provides resource export functionality, generating download links via a web server within the local area network, and supporting access from PCs.

[0094] The access control unit is used to determine the different operation permissions of different users based on role-based access control.

[0095] In one exemplary embodiment, the storage unit specifically includes: SQLite database is used to store structured data.

[0096] A file system is used to store unstructured data.

[0097] The cloud synchronization subunit is used to synchronize different types of data from the intelligent handwriting sample extraction system to the cloud server for data backup and multi-terminal synchronization.

[0098] In practical applications, the data management module is associated with the information collection module: it receives and stores raw collected data (personnel, case, environmental information) and provides a data retrieval interface for other modules to call.

[0099] Associated with the handwriting generation module: Stores the generated results (text, voice files) and manages dictionary / lexicon data (such as local dictionary updates).

[0100] Associated with the sample verification module: Stores verification records (sample file, signature, timestamp), supporting historical queries and audit trails.

[0101] Data management methods: Storage mechanism: Structured data is stored in a local SQLite database, while unstructured data (images, audio, signature files) is stored in a file system (such as Android's internal storage). Cloud synchronization is achieved via HTTPS API.

[0102] Search function: Supports multi-condition search (such as case number, time range), and returns case details (including all related data).

[0103] Resource Management: Provides a resource export function, which generates download links within the local area network via a web server (such as AndServer) and supports access from PCs.

[0104] Access control: Based on role-based access control (RBAC), different users (such as administrators and operators) have different operation permissions (such as view only or modify).

[0105] Algorithm Research: The focus is on text generation and speech synthesis algorithms. The cloud-based text generation component uses a Transformer-based NLP model to process input keywords and generate context-appropriate handwritten text. Speech generation employs a TTS algorithm to convert text into speech and provide voice feedback.

[0106] Data security and traceability: The collection and management of sample data involves highly sensitive evidence. The design incorporates encrypted storage, access control and timestamp recording to ensure that the time and operation process of sample files are traceable throughout their entire lifecycle to prevent data tampering.

[0107] Overall workflow: After the system starts, users first enter the information collection stage through login authentication. The system uses hardware devices such as ID card readers, fingerprint sensors, and signature pads to automatically collect and extract detailed information about the person and the writer. At the same time, case-related data is entered through the user interface. Environmental sensors automatically record environmental information such as GPS coordinates and timestamps. All collected data is transmitted to the handwriting generation module after being structured and is temporarily backed up in the data management module.

[0108] After entering the handwriting generation stage, the system first extracts keywords from the collected data. Users can choose local or cloud generation modes. The system calls the corresponding NLP model to generate handwriting text that meets the requirements based on the keywords. Then, the text is converted into an audio file through the TTS speech synthesis engine. The generated text and audio content are transmitted to the sample confirmation module after quality checks.

[0109] The sample verification stage is a crucial step in ensuring the authenticity and integrity of the samples. Users upload handwriting sample images collected on-site via camera. The system automatically performs keyword matching checks and then requires the writer to provide an electronic signature and fingerprint verification. Biometric technology is used to confirm the authenticity of the identity. At the same time, the system automatically generates a timestamp and binds it to the sample file. All verification data is encrypted before being transmitted to the data management module.

[0110] The data management phase, as the final stage of the system, is responsible for the unified storage and management of all collected information, generated content, and confirmation records. It uses an SQLite database to store structured data and a file system to manage unstructured files, while also supporting cloud synchronization and access control. Users can search historical records or export resource files within the local area network using the search function. After the entire process is completed, the system automatically generates a PDF report containing sample summaries and operation records for users to use for subsequent analysis and archiving.

[0111] This complete workflow not only achieves fully automated processing from information collection to data management, but also ensures the standardization, efficiency and traceability of sample extraction through close cooperation between various modules.

[0112] By standardizing the extraction process and optimizing the extraction methods, the handwriting samples output by the handwriting generation module are clearer and more accurate. This also improves work efficiency, reduces previously cumbersome procedures, and optimizes the user experience. Through an intuitive interface design and simplified operation, extraction personnel can easily and quickly collect handwriting samples, lowering the barrier to entry and enhancing the software's user experience.

[0113] The main achievements of this application will reach an advanced level in terms of technology. The terminal equipment will be portable and will employ technologies such as identity recognition, fingerprint recognition, and voice recognition to improve the efficiency of information collection and the accuracy and speed of text recognition. The handwriting experimental sample extraction assistant will automatically generate voice text. Based on keywords and other information, the system will automatically generate the most suitable text for collection via the cloud or local machine, ensuring that the collected handwriting samples meet the specifications.

[0114] The implementation of this application is expected to generate significant benefits. First, in identification work, efficient handwriting sample collection will greatly improve work efficiency and reduce labor costs. Second, in terms of industrialization, the results of this application will promote the development of related industrial chains, form products with market competitiveness, and bring economic benefits to related industries.

[0115] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0116] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A smart system for extracting handwriting test samples, characterized in that, include: The information collection module is used to collect relevant data and generate structured data; The relevant data includes personnel information, case information, and environmental data; The handwriting generation module is used to process the structured data and generate handwriting content that meets the identification requirements; the handwriting content includes handwriting text and audio files; The sample verification module is used to verify the integrity of the handwriting content and the authenticity of the extractor and the extracted person, and generate a verification result.

2. The intelligent extraction system for handwriting test samples according to claim 1, characterized in that, The information collection module specifically includes: A personnel information collection unit is used to collect personnel information using integrated hardware devices. These integrated hardware devices include an ID card reader, a fingerprint sensor, a signature pad, and a GPS locator. The ID card reader automatically extracts the ID card information of the person being collected using OCR technology. The fingerprint sensor is a capacitive fingerprint sensor used to collect the biometric features of the person being collected and generate an encrypted fingerprint template. The signature pad is used to capture the signature trajectory of the person being collected and store it as a vector graphic. The GPS locator automatically obtains the coordinates of the collection location. The personnel information includes the ID card information, the encrypted fingerprint template, the vector graphic, and the coordinates of the collection location. The case information collection unit is used to generate a unique extraction number based on a timestamp and a random number according to the input fields, and to collect the case information according to the extraction number; the input fields include case number, client, case name, case type, extraction location, detailed address, and submission time; An environmental data acquisition unit is used to automatically record environmental data using device sensors; the environmental data includes timestamps, device ID, operator ID, GPS coordinates, and network status. The structured data generation unit is used to generate structured data from the personnel information, the case information, and the environmental data.

3. The intelligent handwriting sample extraction system according to claim 1, characterized in that, The handwriting generation module specifically includes: The handwriting text generation unit is used to call an NLP model to preprocess the structured data based on the custom keyword content and generate handwriting text that conforms to the context. An audio file generation unit is used to convert the handwritten text into an audio file; The handwriting content generation unit is used to generate handwriting content based on the handwriting text and the audio file.

4. The intelligent handwriting sample extraction system according to claim 3, characterized in that, The handwriting text generation unit specifically includes: The generation mode selection subunit is used to select the generation mode based on the custom keyword content; the generation modes include local mode and cloud mode. The local mode handwriting text generation subunit is used to invoke a local NLP model and use a pre-trained word-filling and sentence-building expansion model to generate handwriting text that conforms to the context when the generation mode is local mode; the local NLP model is a lightweight model and the lightweight model is the ONNX Runtime framework. The cloud-based handwriting text generation subunit is used to call a Transformer-based cloud NLP model to generate handwriting text that conforms to the context when the generation mode is cloud-based; the cloud NLP model includes a Chinese model and an English model.

5. The intelligent handwriting sample extraction system according to claim 3, characterized in that, The audio file generation unit specifically includes: The post-processing subunit is used to perform post-processing on the handwriting text and determine the processed handwriting text. The speech file generation subunit is used to perform speech synthesis using TTS technology to generate speech files.

6. The intelligent extraction system for handwriting test samples according to claim 1, characterized in that, The sample verification module specifically includes: The sample uploading unit is used to acquire handwriting sample images collected on-site; The keyword matching degree checking unit is used to perform keyword matching degree checking between the handwriting sample image and the handwriting content, determine whether the keyword matching degree is greater than the matching degree threshold, and determine the first judgment result. The integrity verification unit is used to verify the integrity of the handwriting content based on the digital fingerprint of the handwriting sample image. The biometric verification unit is used to perform biometric verification on the extracted person based on the personnel information in the information collection module, determine whether the biometric verification is successful, and confirm the second judgment result. The confirmation result generation unit is used to determine that the authenticity of the extracted person has been verified, record the current timestamp, generate a UTC timestamp, bind the handwriting sample image with the handwriting content, and generate a confirmation result; The return unit is used to return to the information collection module, re-collect relevant data, and generate structured data.

7. The intelligent extraction system for handwriting test samples according to claim 6, characterized in that, The biometric verification unit specifically includes: The integrity verification subunit is used to calculate the digital fingerprint of the handwriting sample image using a hash algorithm, and to verify the integrity of the handwriting content based on the digital fingerprint. The input subunit is used to input the fingerprint information of the person being extracted; The similarity calculation subunit is used to calculate the similarity between the fingerprint information and the fingerprint template in the information acquisition module using the Hamming distance algorithm; The authenticity verification subunit is used to determine that the authenticity verification of the extracted person has passed when the similarity is greater than a set threshold.

8. The intelligent extraction system for handwriting test samples according to claim 1, characterized in that, Also includes: The data management module is used to manage the data storage, retrieval, backup, and security management of the intelligent extraction system for handwriting experimental samples.

9. The intelligent extraction system for handwriting test samples according to claim 8, characterized in that, The data management module specifically includes: The storage unit is used to store different types of data from the intelligent handwriting sample extraction system and to synchronize different types of data to the cloud server. The retrieval unit supports multi-condition retrieval and returns case details; the multi-condition retrieval includes case number retrieval and time range retrieval; the case details include all related data related to the case. The resource management unit provides resource export functionality, generating download links via a web server within the local area network, and supporting PC access. The access control unit is used to determine the different operation permissions of different users based on role-based access control.

10. The intelligent extraction system for handwriting test samples according to claim 9, characterized in that, The storage unit specifically includes: SQLite database, used to store structured data; A file system is used to store unstructured data; The cloud synchronization subunit is used to synchronize different types of data from the intelligent handwriting sample extraction system to the cloud server for data backup and multi-terminal synchronization.