Off-line double-recording real-time verification method and device
By employing an offline dual-recording real-time verification method, preloading resources, and constructing a dialect-adaptive dictionary, combined with localized face recognition, the problem of recording interruption and recognition failure in the insurance dual-recording system under weak network conditions was solved. This enabled continuous recording and real-time quality inspection in offline environments, improving the system's fault tolerance and compliance verification efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PICC INFORMATION TECH CO LTD
- Filing Date
- 2025-12-16
- Publication Date
- 2026-05-01
AI Technical Summary
Existing insurance dual recording systems suffer from high interruption rates in weak network environments, high voice recognition failure rates, and an inability to detect compliance risks in real time, resulting in low business efficiency and increased labor costs.
An offline dual-recording real-time verification method is adopted. By preloading ID card images, facial feature data and speech recognition model resources, a dialect-adaptive word library is constructed. The network signal is monitored in real time and a localized facial recognition algorithm is used to achieve recording and quality inspection in environments with no network or weak network.
To achieve continuous recording and real-time quality inspection of the insurance dual recording process in environments with no or weak network coverage, reduce the impact of network fluctuations, improve system fault tolerance and compliance verification efficiency, and reduce the workload of manual quality inspection.
Smart Images

Figure CN121967775A_ABST
Abstract
Description
A method and apparatus for offline dual-recording real-time verification Technical Field
[0001] This invention belongs to the field of data processing technology, specifically relating to an offline dual-recording real-time verification method and device. Background Technology
[0002] Dual recording in insurance, as a core technology for compliance management in the insurance industry, is widely used in the audio and video recording and quality inspection stages of the sales process. Among related technologies, a complete technical system for dual recording has been constructed through the collaborative operation of voice broadcasting, voice recognition, and facial recognition. Specifically, this technology covers the entire process from document display and identity verification to terms confirmation, including key aspects such as real-time voice content quality inspection, face detection within the same frame, and signature verification. With the rapid development of mobile insurance business, traditional dual recording systems have gradually exposed their strong dependence on the network environment, especially in remote areas or complex scenarios, where network fluctuations leading to business interruptions are becoming increasingly prominent.
[0003] However, existing dual-recording methods directly utilize cloud-based voice broadcasting and recognition technologies without establishing a localized offline processing mechanism, potentially leading to a recording process interruption rate as high as 30%. Specifically, existing technologies typically rely on real-time network transmission of document images and voice models, but suffer from limitations such as document display delays in weak network environments and voice recognition failure rates exceeding 40%. Dialectal accents are particularly problematic, as traditional voice recognition models cannot effectively handle regional speech characteristics, resulting in compliance verification failures. Consequently, manual quality inspection requires post-recording checks, which cannot detect compliance risks such as missing personnel or misleading answers in real time, thus impacting business processing efficiency and increasing labor costs. Industry data shows that manual quality inspection accounts for over 70% of the dual-recording process, and inspection delays often exceed minutes, failing to meet real-time requirements. Summary of the Invention
[0004] The present invention aims to at least partially solve one of the technical problems in the related art.
[0005] Therefore, the first objective of this invention is to propose an offline dual-recording real-time verification method.
[0006] The main objective of this invention is to realize an intelligent dual-recording system by combining technologies such as real-time network monitoring, resource preloading, offline voice broadcasting, offline voice recognition, and offline face detection. This system greatly reduces the impact of network fluctuations during the recording process and significantly improves quality inspection efficiency through face detection and voice recognition.
[0007] The second objective of this invention is to provide an offline dual-recording real-time verification device.
[0008] The third objective of this invention is to provide a computer device.
[0009] A fourth objective of this invention is to provide a non-transitory computer-readable storage medium.
[0010] To achieve the above objectives, the first aspect of this invention proposes an offline dual-recording real-time verification method, comprising: S1, preloading and caching document images, facial feature data, and speech recognition model resources, and implementing incremental updates through a version control mechanism; S2, constructing a dialect-adapted word library, and dynamically expanding the homophone matching library based on word frequency analysis of salesperson responses and manual quality inspection feedback data; S3, monitoring the network signal strength during the recording process in real time, triggering a visual warning and prompting a network switch when a persistent weak network state is detected; S4, using a locally deployed facial recognition algorithm to perform real-time same-frame detection on the recorded screen, judging the consistency of personnel identities based on a preset similarity threshold, and generating a verification result.
[0011] In one embodiment of the present invention, S1 includes: S11, generating a resource key value based on the user ID, resource URL, token, and timestamp, and calculating... Implement resource version control; S12, when a new resource's key value is detected to be inconsistent with the local cache, download the base64 format image resource first, convert it to an encrypted storage format, and then update the local cache.
[0012] In one embodiment of the present invention, S2 includes: S21, when the frequency of a word in the answer to the same question in the region exceeds 70%, the word is marked as an item to be expanded and added to the word library; S22, if the manual quality inspection platform marks the answer in the video as a homophone of the keyword in the word library, the answer is directly added to the word library as a homophone.
[0013] In one embodiment of the present invention, step S4 further includes: S41, extracting facial feature values after liveness detection based on the Megvii SDK. , with pre-stored feature values Compare the similarities. If the same person is identified at the time; S42, based on the results of the same-frame detection feedback from manual quality inspection, the proportion of excessively high, excessively low and accurate detection thresholds is statistically analyzed weekly, and the similarity threshold benchmarks of different institutions are dynamically adjusted.
[0014] In one embodiment of the present invention, the method further includes: S5, performing an encryption storage operation on the preloaded document image, and generating an encrypted file using the AES-256 algorithm. During the recording process, a decryption algorithm is used. Enables fast loading and display in offline environments.
[0015] To achieve the above objectives, a second aspect of the present invention proposes an offline dual-recording real-time verification device, comprising: a resource preloading and caching module for preloading and caching document images, facial feature data, and speech recognition model resources, and implementing incremental updates through a version control mechanism; a dialect adaptation dictionary construction module for constructing a dialect adaptation dictionary, dynamically expanding the homophone matching library based on the word frequency analysis of salesperson answers and manual quality inspection feedback data; a network signal monitoring and early warning module for real-time monitoring of network signal strength during the recording process, triggering a visual early warning and prompting a network switch when a continuous weak network state is detected; and a face recognition co-op detection module for using a locally deployed face recognition algorithm to perform real-time co-op detection on the recorded screen, judging the consistency of personnel identities based on a preset similarity threshold, and generating a verification result.
[0016] The present invention discloses an offline dual-recording real-time verification method and apparatus, which can realize continuous recording and real-time quality inspection of the insurance dual-recording process in environments without network or with weak network, effectively reducing the impact of network fluctuations on voice broadcasting, voice recognition and face verification, and improving system fault tolerance and compliance verification efficiency.
[0017] To achieve the above objectives, a third aspect of this application provides a computer device, including a processor and a memory; wherein the processor reads executable program code stored in the memory to run a program corresponding to the executable program code, for implementing an offline dual-recording real-time verification method as described in the first aspect embodiment.
[0018] To achieve the above objectives, the fourth aspect of this application provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements an offline dual-recording real-time verification method as described in the first aspect embodiment.
[0019] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0020] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of embodiments taken in conjunction with the accompanying drawings, in which: FIG1 is a flowchart of an offline dual-recording real-time verification method according to an embodiment of the present invention; FIG2 is a resource update flowchart according to an embodiment of the present invention; FIG3 is a flowchart of a recognition lexicon construction flowchart according to an embodiment of the present invention; FIG4 is a face verification processing flowchart according to an embodiment of the present invention; FIG5 is a structural diagram of an offline dual-recording real-time verification device according to an embodiment of the present invention; FIG6 is a computer device according to an embodiment of the present invention. Detailed Implementation
[0021] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0022] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0023] The following description, with reference to the accompanying drawings, describes an offline dual-recording real-time verification method and apparatus according to an embodiment of the present invention.
[0024] Example 1 Figure 1 is a flowchart of an offline dual-recording real-time verification method according to an embodiment of the present invention. As shown in Figure 1, it includes: S1, preloading and caching document images, facial feature data and speech recognition model resources, and implementing incremental updates through a version control mechanism.
[0025] In some implementations, this step first uses a resource synchronization mechanism between the server and client to download and cache all required document images, face images, and speech recognition model resources to local storage. The document images are cached using encrypted storage to ensure data security. The face data caching involves two key stages: First, the face object `face1` is extracted from the original document image using the `faceDetect` algorithm. Then, liveness detection is performed to generate a live face object `face2`. The system compares the feature vectors of `face1` and `face2`. If the similarity exceeds 80%, they are determined to be the same person, and this `face2` data will be used for real-time face comparison during the recording process.
[0026] Specifically, the facial similarity threshold is set to 80%, which can be dynamically fine-tuned based on actual quality inspection results. The key value for caching ID card images is generated by processing the user ID, resource URL, token, and timestamp using MD5, ensuring resource version traceability. If the resource changes, the key value is updated accordingly, allowing the client to determine whether to re-download the resource, thus achieving incremental updates, avoiding duplicate loading, and improving system response speed and resource utilization.
[0027] Further, S1 includes: S11, generating a resource key value based on the user ID, resource URL, token, and timestamp, and calculating... Implement resource version control.
[0028] In some implementations, this key value is generated as follows: The key consists of four parameters: user ID (a unique string identifying the user), URL (the resource's access address or Base64-encoded image data), token (a temporary access token used for authentication), and timestamp (the current system timestamp, usually in milliseconds). In practice, the system first concatenates these four parameters into a string in a fixed order, then calculates its hash value using the MD5 algorithm, generating a 128-bit hexadecimal string as the resource's unique key. This key identifies the current version status of the resource; any change in any parameter will alter the key, triggering a re-download or update of the resource.
[0029] Specifically, the user ID is typically a string between 10 and 30 characters long, the URL can be a standard HTTP address or Base64 encoded binary data, the token is a 64-bit JWT format token, and the timestamp is a 13-bit millisecond-level timestamp. The MD5 algorithm outputs a fixed-length 32-bit hexadecimal string, which has good anti-collision properties and computational efficiency, making it suitable for resource version control scenarios.
[0030] S12, when a new resource's key value is detected to be inconsistent with the local cache, the base64 format image resource is downloaded first and converted to an encrypted storage format before updating the local cache.
[0031] In some implementations, when a new resource's key value is detected to be inconsistent with the local cache, the system will prioritize downloading the base64 format image resource and converting it to an encrypted storage format to ensure the accuracy and security of resource updates. This step is technically based on a resource version control mechanism and a local cache consistency verification strategy, determining whether a resource has changed by comparing key values. The key value is generated by processing the user ID, resource URL, token, and timestamp using the MD5 algorithm, ensuring its uniqueness and immutability, thus providing a reliable basis for resource version management.
[0032] Specifically, the system compares the key value sent by the server with the key value in the local cache. If they do not match, a resource update process is triggered. Base64 format image resources have good compatibility due to their encoding method, making them suitable for fast transmission and parsing in weak or no network environments. After the image is downloaded, the system uses symmetric encryption algorithms such as AES-256 or SM4 to encrypt the plaintext base64 data, converting it into encrypted binary format and storing it as a file in a local secure directory, ensuring the confidentiality and integrity of the data during local storage.
[0033] S2, construct a dialect-adapted word library, and dynamically expand the homophone matching library based on the word frequency analysis of salespersons' answers and the feedback data from manual quality inspection.
[0034] In some implementations, the construction of this dictionary is divided into two stages: initial dictionary construction and dynamic expansion. The initial dictionary is customized by each branch office according to actual business needs, containing keywords to be recognized and their corresponding pinyin representations. For example, if a business process requires a customer to answer "I have read and understand the insurance terms," the dictionary will contain keywords such as "read," "understand," and "insurance terms" along with their pinyin sequences. This stage ensures that the system has basic recognition capabilities under standard voice input.
[0035] Furthermore, the dynamic expansion of the vocabulary relies on two data sources: an automatic learning mechanism from the app and manual feedback from the quality inspection platform. On the app, the system analyzes customer responses in real time using a speech recognition module, calculating the recognition frequency of each keyword. When a non-standard answer has a frequency exceeding 70% and sounds like a standard keyword, the system marks that word as an item to be expanded and automatically adds it to the vocabulary. For example, if a customer repeatedly misreads "understand" as "solve," and the recognition frequency exceeds a threshold, "solve" will be added to the homophone matching database to improve the recognition success rate.
[0036] On the quality inspection platform, when reviewing recorded videos, quality inspectors can manually mark the answer as "needs expansion" if they find that the system failed to recognize the keywords in the customer's actual response. The system will then add these homophones or near-homophones to the dictionary based on the quality inspection feedback data, thereby achieving collaborative optimization between manual and automated processes.
[0037] Furthermore, the dictionary also features an optimization mechanism. For dynamically added keywords, if their pinyin or pronunciation features already exist in the dictionary, only their word frequency count is increased. The system periodically performs statistical analysis on the dictionary, removing low-frequency words with frequencies below a set threshold to ensure the dictionary content has regional universality and representativeness. This mechanism effectively prevents dictionary expansion and improves recognition efficiency.
[0038] Furthermore, S2 includes: S21, when the frequency of a word in the answer to the same question exceeds 70% in the region, the word is marked as an item to be expanded and added to the lexicon.
[0039] From a technical implementation perspective, this step first uses a speech recognition engine to convert the user's answers to specific questions during recording into text. Then, keyword matching and word frequency analysis are performed within a preset semantic region (such as the response range for a certain type of question). Specifically, the system maintains a dynamically updated set of keywords. Each keyword Each response corresponds to a standard word or phrase. During the recognition process, the system will use the user's actual response... With each of the keyword sets Perform similarity comparison of pinyin or acoustic features, if With a certain If the pinyin matching accuracy exceeds a preset threshold (e.g., 85%), then it is considered... yes Homophone variants. The system further statistically analyzes all identified homophones within this region. China and a certain The percentage of homophones in frequency; if this percentage exceeds 70%, then... Marked as items to be expanded, and added to the thesaurus via the backend interface. This is to enhance the model's ability to recognize regional accents and semantic variations.
[0040] At the parameter level, this step involves several key parameters, including the pinyin matching threshold. Word frequency statistics window size Frequency percentage threshold .in, Typically set to To ensure the accuracy of homophone recognition; It can be set according to the actual business scenario. to The results of this identification are used to statistically analyze high-frequency words in the region.
[0041] S22. If the manual quality inspection platform marks the response in the video as a homophone of a keyword in the dictionary, then the response will be directly added to the dictionary as a homophone.
[0042] In some implementations, the system first converts the audio content of the recorded video into text using an offline speech recognition engine, and then compares the recognition result with keywords in a dictionary at the pinyin level. Specifically, the system uses a pinyin-based speech similarity algorithm to match the recognized response with the pinyin sequence of keywords in the dictionary character by character. If the recognized text and keywords are completely identical at the pinyin level, but the Chinese characters differ (i.e., homophones), the system determines it as a homophone match. This determination logic can be expressed as: if and This triggers the homophone addition mechanism.
[0043] Specifically, the key parameters involved in this step include the pinyin matching threshold, the word database update frequency, and the criteria for adding homophones. The pinyin matching threshold is usually set to 100%, requiring complete consistency; the word database update frequency can be set to daily or automatically updated after each recorded conversation; the criteria for adding homophones is the labeling result of the manual quality inspection platform, that is, the quality inspector confirms that the answer is compliant content in the video, but it is misjudged as a mismatch because the speech recognition model fails to correctly recognize the Chinese characters.
[0044] S3 monitors the network signal strength in real time during the recording process. When a persistent weak network condition is detected, it triggers a visual warning and prompts the user to switch networks.
[0045] In some implementations, this step involves continuously monitoring the network status during the recording process of the Android dual-recording app. System-level network APIs (such as Android's ConnectivityManager) are used to obtain metrics such as the Received Signal Strength Indicator (RSSI) value or Round-Trip Time (RTT) of the current network connection to quantitatively evaluate network quality. In specific implementations, the system sets a sliding time window (e.g., a 10-second window) and averages the network signal strength within the window. If the average RSSI value is lower than a set weak network threshold (e.g., -90 dBm) or the RTT exceeds 500 ms, the network is considered weak.
[0046] Furthermore, the system employs a state machine mechanism to continuously assess the weak network status. Only when the weak network status is triggered within multiple consecutive window periods (e.g., three consecutive windows) is it considered a "persistent weak network," thus avoiding false alarms caused by instantaneous network fluctuations. Once a persistent weak network is confirmed, the system will trigger a visual warning through the UI layer. For example, a semi-transparent warning box will pop up on the recording interface, displaying prompts such as "The current network signal is weak. It is recommended to switch to Wi-Fi or 4G / 5G network to ensure recording quality," accompanied by a signal strength icon (e.g., 1-5 signal bars) dynamically displaying the current network quality.
[0047] S4 uses a locally deployed face recognition algorithm to perform real-time same-frame detection on the recorded screen, judges the consistency of the person's identity based on a preset similarity threshold, and generates a verification result.
[0048] In some implementations, this step employs a locally deployed face recognition algorithm to perform real-time frame detection on the recorded footage. This technology is based on a deep learning-based face feature extraction and comparison mechanism. Specifically, before recording, the system preprocesses the participants' local face images using the `faceDetect` module, extracting the feature vectors of the face objects (`face2`) and storing them in a local database. During recording, the system captures video frames in real-time via a camera, uses face detection algorithms (such as MTCNN or RetinaFace) to locate faces in each frame, extracts the feature vectors of faces in the current frame, and compares their similarity with the pre-stored `face2` features.
[0049] This face recognition algorithm employs a feature encoding model based on deep convolutional neural networks (CNNs), such as FaceNet or ArcFace, to map face images into a high-dimensional feature space and perform matching using Euclidean distance or cosine similarity. In this system, a similarity threshold of 70% is set. When the similarity between a real-time detected face and pre-stored features exceeds this threshold, it is determined to be the same person; otherwise, it is considered an inconsistency. This threshold can be dynamically fine-tuned based on quality inspection feedback from different institutions to adapt to the recognition accuracy requirements in different scenarios.
[0050] Furthermore, the system continuously performs face detection during recording to ensure that all personnel who should be on camera are always in the frame. If a person is detected as missing, a pop-up notification will appear in the UI to alert the salesperson in real time, preventing recording interruptions or the generation of non-compliant content. This step can still operate stably in environments without or with weak network connectivity, greatly improving the robustness and automation of the dual-recording process.
[0051] Furthermore, S4 includes: S41, extracting facial feature values after liveness detection based on the Megvii SDK. , with pre-stored feature values Compare the similarities. They were determined to be the same person.
[0052] In some implementations, this step involves extracting and comparing facial feature values after liveness detection using the Megvii SDK to achieve real-time identity verification of participants during offline dual recording. Specifically, the system first extracts static facial feature values of the recording personnel through the face detection module (faceDetect) in the pre-preparation stage. Subsequently, during the recording process, video frames were captured in real time via a camera, and the liveness detection function of the Megvii SDK was used to determine whether the captured face was a real, living person. Liveness detection employs techniques such as multi-frame image analysis, micro-expression recognition, blink detection, and head motion recognition to ensure the accuracy of the extracted facial feature values. It is based on real human faces, not fakes such as photos, videos, or 3D models.
[0053] Furthermore, the system inputs the face image after liveness detection into the face feature extraction module of the Megvii SDK to generate the corresponding feature vector. This feature vector is typically a 128-dimensional or higher-dimensional floating-point array used to represent the biometric information of a human face. Subsequently, the system will... With pre-stored Perform cosine similarity calculation to obtain the similarity value. .when When the system determines that the face in the current video frame is the same person as the pre-stored face, it completes the identity verification.
[0054] S42. Based on the results of same-frame detection feedback from manual quality inspection, the proportion of detections with excessively high or low thresholds and accurate detections is statistically analyzed weekly, and the similarity threshold benchmarks of different institutions are dynamically adjusted.
[0055] In some implementations, the system first performs frame-by-frame face detection and comparison on the received dual-recorded video in a manual quality inspection platform to determine whether the current frame contains the person who should be on camera. Based on system prompts, quality inspectors categorize the detection results into three types: too high threshold (i.e., actually in the shot but not recognized by the system), too low threshold (i.e., not actually in the shot but misidentified by the system), and accurate detection (i.e., the recognition result matches the actual person). These categorized results are synchronized to the backend data processing module via the quality inspection platform interface. The system then calculates the percentage of each of the three types of detection results weekly based on the categorized results.
[0056] in, , , These represent the number of samples detected when the threshold is too high, too low, or precisely. This represents the total number of samples. The system calculates based on... and The distribution of these factors is analyzed, and a weighted average strategy is used to fine-tune the similarity threshold benchmark for the institutions. For example:
[0057] in, The initial threshold for the institution, The adjustment coefficient is usually set to... This is to ensure the smoothness and stability of threshold adjustment.
[0058] S5 performs an encryption storage operation on the preloaded document image, generating an encrypted file using the AES-256 algorithm. During the recording process, a decryption algorithm is used. Enables fast loading and display in offline environments.
[0059] In some implementations, this step uses the AES-256 (Advanced Encryption Standard, 256-bit key length) algorithm to encrypt the document image, generating an encrypted file. This ensures that document information is not illegally accessed or tampered with during local caching. AES-256 is a symmetric encryption algorithm whose encryption process is based on a 14-round Feistel structure. Each round includes operations such as byte substitution (SubBytes), row shifting (ShiftRows), column mixing (MixColumns), and round key addition (AddRoundKey). It has high security and computational efficiency and complies with the FIPS 197 standard set by NIST (National Institute of Standards and Technology).
[0060] In its implementation, the system first standardizes and compresses the ID image using JPEG or PNG format, typically setting the compression quality parameter to 85% to balance image clarity and storage efficiency. Then, it uses a preset encryption key... Image data is encrypted using a 256-bit key and CBC (Cipher Block Chaining) mode to enhance encryption strength. The encrypted file... It is stored in binary form in the local file system, and the path is composed of user ID, document type and timestamp to ensure uniqueness and traceability.
[0061] During the recording process, the system calls the decryption algorithm. This technology enables rapid loading and display of ID card images even in offline environments. The decryption algorithm also adheres to the AES-256 standard, using the same key and mode as the encryption process to ensure data consistency. This step is used in the ID card display stage of the dual-recording process, supporting rapid retrieval and decryption of ID card images under weak or no network conditions, avoiding recording interruptions caused by network latency or outages.
[0062] This invention provides an offline dual-recording real-time verification method that enables continuous recording and real-time quality inspection of the insurance dual-recording process in environments with no or weak network coverage. This effectively reduces the impact of network fluctuations on voice broadcasting, voice recognition, and face verification, thereby improving system fault tolerance and compliance verification efficiency.
[0063] Example 2: The following describes in detail an offline dual-recording real-time verification method according to an embodiment of the present invention, with reference to the accompanying drawings.
[0064] S10. Data preparation before recording.
[0065] This step implements a data caching mechanism, including caching of recorded identification documents, caching of recorded facial data, caching of offline playback during recording, and caching of initialization resources for offline recognition.
[0066] S101. Local image resource caching and updating are shown in Figure 2. The dual-recording server obtains the recording participants' information and corresponding identification information from the integrated system. Each resource corresponds to a unique key value; if the resource changes, the key value will also change. The dual-recording app obtains resource information from the server, queries the local cache based on the download address and key value, and re-downloads the resource if it has changed. If there are changes to image resources in the app, the image is first uploaded to the private cloud. After successful upload, the cloud address is notified to the backend, and the backend updates the resource key and URL.
[0067] Key generation rules: User ID + URL + token + timestamp (processed MD5 value); URL has two values: a normal HTTP link and direct image base64 data, distinguished by type.
[0068] S102. Caching of Face Data. After the above processing, the local face image of each person being recorded is obtained. The face image is then processed using `faceDetect` to extract the corresponding face object `face1`. Liveness detection is then performed. If the liveness detection is successful, the live face `face2` is extracted. The feature values of `face1` and `face2` are compared. If the similarity exceeds 80%, they are considered to be the same person. This `face2` data is retained for real-time face comparison during the recording process.
[0069] S103. Loading offline voice resources. Resources required for offline functionality are pre-stored on the server, and downloaded and loaded into the app during initialization.
[0070] S20, Intelligent guidance scheme during the recording process.
[0071] S201. Continuous Network Signal Monitoring: Throughout the recording process, monitor the network for each data stream. Display the recorder's ID and a signal icon representing network strength on each window. If the network condition remains poor, a pop-up window should remind the salesperson to pay attention and switch networks to avoid affecting the recording process and business operations.
[0072] S202, Offline Broadcasting, Recognition, and Recognition Matching Library Construction: This section addresses the issues of text broadcasting and speech recognition in offline networks through machine learning and model algorithms. Due to the limitations of offline models, the recognized text and keyword matching must be precisely matched to pass. In reality, since sales staff are distributed across the country of various ages, accent issues can cause recognition failures and automatic skipping to the next step, requiring manual confirmation to proceed. To address this issue, a keyword library is specifically created to maximize compatibility with accent issues and recognize homophones. The recognition dictionary construction process is shown in Figure 3.
[0073] a. Initial vocabulary: This includes keywords to be identified and their corresponding pinyin. Keywords are set by each branch office, and pinyin is added automatically by the system.
[0074] b. Thesaurus expansion: There are two sources. The first part comes from the app. The app judges the matching of words with keywords based on the frequency ratio and homophones of the words in the answers. If the word ratio of the answer to the same question exceeds 70%, it is marked as needing expansion. Answer words that are homophones of keywords are directly added to the thesaurus. The second part comes from the quality inspection platform. It is judged manually based on the answers in the video. If the corresponding answer is marked as needing expansion, it is added to the thesaurus.
[0075] c. Keyword optimization. For keywords dynamically added during recording, if they already exist in the database, their frequency is incremented by 1. Frequency statistics are performed regularly, and low-frequency keywords are cleaned up to make them more universally applicable within the region.
[0076] S30, Page preloading.
[0077] During recording, identification information needs to be displayed, typically as images. To ensure fast display during recording, the encrypted identification data is pre-cached locally during the preparation phase. When loading the identification, simply informing the web page of the local identification address enables offline display.
[0078] S40, full-process face verification.
[0079] Face recognition and comparison: This system uses Megvii's face recognition technology to compare the similarity of faces in all preview cameras and lenses. If the comparison value is greater than the threshold, it is considered to be the same person. The initial threshold can be continuously fine-tuned based on the quality inspection results.
[0080] Verification reminder during recording: Based on all face data from the previous cached process, full-process face detection and in-frame prompts are performed during recording, displaying in real time whether any personnel are missing from the recording. The processing flow is shown in Figure 4.
[0081] Similarity threshold fine-tuning: In conjunction with the existing manual quality inspection system, the accuracy of in-frame results is categorized. Results that are actually in the frame but are flagged as not in the shot are marked as having a high threshold; results that are not in the frame but are not flagged are marked as having a low threshold; and results that are correct are marked as accurate. Different initial thresholds are set for different institutions, and the accuracy of different thresholds is compared. The proportion of the three types of results is statistically analyzed weekly, and the comparison thresholds are dynamically fine-tuned as needed. Ultimately, the optimal threshold benchmark is found.
[0082] By implementing the above methods, the audio and video recording process is largely unaffected by network fluctuations. The entire process is automated, requiring only viewing the broadcast content and answering system prompts, eliminating reliance on manual clicks. AI-powered alerts help users identify non-compliant behavior, reducing communication costs. Various verification measures during recording ensure improved real-time quality control efficiency and reduce the workload of manual quality checks.
[0083] Example 3 To implement the above embodiments, as shown in Figure 5, this embodiment also provides an offline dual-recording real-time verification device 10. The device 10 includes a resource preloading and caching module 100, a dialect adaptation dictionary construction module 200, a network signal monitoring and early warning module 300, and a face recognition in-frame detection module 400.
[0084] The resource preloading and caching module 100 is used to preload and cache ID images, facial feature data, and speech recognition model resources, and implements incremental updates through a version control mechanism; the dialect adaptation dictionary construction module 200 is used to build a dialect adaptation dictionary, and dynamically expands the homophone matching library based on the word frequency analysis of salespersons' answers and manual quality inspection feedback data; the network signal monitoring and early warning module 300 is used to monitor the network signal strength in real time during the recording process, and triggers a visual early warning and prompts to switch networks when a continuous weak network state is detected; the face recognition co-op detection module 400 is used to perform real-time co-op detection on the recorded screen using a locally deployed face recognition algorithm, judge the consistency of personnel identities based on a preset similarity threshold, and generate verification results.
[0085] Furthermore, the aforementioned resource preloading and caching module 100 is also used to: generate a resource key value based on the user ID, resource URL, token, and timestamp, and calculate... Implement resource version control; when a new resource's key value is detected to be inconsistent with the local cache, download the base64 format image resource first, convert it to an encrypted storage format, and then update the local cache.
[0086] Furthermore, the aforementioned dialect adaptation dictionary construction module 200 is also used to: mark the word as an item to be expanded and add it to the dictionary when the frequency of the word in the answer to the same question exceeds 70% in the region; if the manual quality inspection platform marks the answer in the video as a homophone but different character from the keyword in the dictionary, then the answer is directly added to the dictionary as a homophone.
[0087] Furthermore, the aforementioned face recognition co-frame detection module 400 is also used to: extract facial feature values after liveness detection based on the Megvii SDK. , with pre-stored feature values Compare the similarities. If the same person is identified at the time, the similarity threshold is determined to be the same person based on the results of the same-frame detection provided by manual quality inspection. The proportion of detections with excessively high or low thresholds and accurate detections is statistically analyzed weekly, and the similarity threshold benchmarks of different institutions are dynamically adjusted.
[0088] Furthermore, device 10 also includes: an encrypted storage module, which performs encrypted storage operations on the pre-loaded document image and generates an encrypted file using the AES-256 algorithm. During the recording process, a decryption algorithm is used. Enables fast loading and display in offline environments.
[0089] An offline dual-recording real-time verification device according to an embodiment of the present invention can realize continuous recording and real-time quality inspection of the insurance dual-recording process in environments without network or with weak network, effectively reducing the impact of network fluctuations on voice broadcasting, voice recognition and face verification, and improving system fault tolerance and compliance verification efficiency.
[0090] To implement the methods of the above embodiments, the present invention also provides a computer device, as shown in FIG6. The computer device 600 includes a memory 601 and a processor 602; wherein the processor 602 reads executable program code stored in the memory 601 to run a program corresponding to the executable program code, so as to implement the various steps of the methods described above.
[0091] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing embodiments.
[0092] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0093] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
Claims
1. An offline dual-recording real-time verification method, characterized in that, include: S1: Preload and cache ID images, facial feature data, and speech recognition model resources, and implement incremental updates through a version control mechanism; S2: Build a dialect-adapted word library, and dynamically expand the homophone matching library based on the word frequency analysis of salespersons' answers and manual quality inspection feedback data; S3: Monitor the network signal strength in real time during the recording process, and trigger a visual warning and prompt to switch networks when a continuous weak network state is detected; S4: Use a locally deployed facial recognition algorithm to perform real-time frame detection on the recorded screen, judge the consistency of personnel identities based on a preset similarity threshold, and generate verification results.
2. The method as described in claim 1, characterized in that, S1 includes: S11, generating a resource key value based on the user ID, resource URL, token, and timestamp, and calculating... Implement resource version control; S12, when a new resource's key value is detected to be inconsistent with the local cache, download the base64 format image resource first, convert it to an encrypted storage format, and then update the local cache.
3. The method as described in claim 1, characterized in that, S2 includes: S21, when the frequency of a word in the answer to the same question exceeds 70% in the region, the word is marked as an item to be expanded and added to the word library; S22, if the manual quality inspection platform marks the answer in the video as a homophone of the keyword in the word library, the answer is directly added to the word library as a homophone.
4. The method as described in claim 1, characterized in that, S4 further includes: S41, extracting facial feature values after liveness detection based on the Megvii SDK. , with pre-stored feature values Compare the similarities. If the same person is identified at the time; S42, based on the results of the same-frame detection feedback from manual quality inspection, the proportion of excessively high, excessively low and accurate detection thresholds is statistically analyzed weekly, and the similarity threshold benchmarks of different institutions are dynamically adjusted.
5. The method as described in claim 1, characterized in that, Also includes: S5 performs an encryption storage operation on the preloaded document image, generating an encrypted file using the AES-256 algorithm. During the recording process, a decryption algorithm is used. Enables fast loading and display in offline environments.
6. An offline dual-recording real-time verification device, characterized in that, include: The resource preloading and caching module is used to preload and cache ID images, facial feature data, and speech recognition model resources, and implements incremental updates through a version control mechanism; the dialect adaptation dictionary construction module is used to build a dialect adaptation dictionary, and dynamically expands the homophone matching library based on the word frequency analysis of salespersons' answers and manual quality inspection feedback data; the network signal monitoring and early warning module is used to monitor the network signal strength in real time during the recording process, and triggers a visual early warning and prompts to switch networks when a continuous weak network state is detected. The face recognition co-op detection module is used to perform real-time co-op detection on the recorded screen using a locally deployed face recognition algorithm, and to determine the consistency of the person's identity based on a preset similarity threshold and generate a verification result.
7. The apparatus as claimed in claim 6, characterized in that, The resource preloading and caching module is also used to: generate a resource key value based on the user ID, resource URL, token, and timestamp, and calculate... Implement resource version control; when a new resource's key value is detected to be inconsistent with the local cache, download the base64 format image resource first, convert it to an encrypted storage format, and then update the local cache.
8. The apparatus as claimed in claim 6, characterized in that, The dialect-adaptive lexicon construction module is also used to: mark a word as an item to be expanded and add it to the lexicon when the frequency of the word in the answer to the same question exceeds 70% in the region; if the manual quality inspection platform marks the answer in the video as a homophone of the keyword in the lexicon, then the answer is directly added to the lexicon as a homophone.
9. A computer device, characterized in that, It includes a processor and a memory; wherein the processor runs a program corresponding to the executable program code stored in the memory to implement an offline dual-recording real-time verification method as described in any one of claims 1-5.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements an offline dual-recording real-time verification method as described in any one of claims 1-5.