Special population assisted communication system and method thereof
Patent Information
- Application Number
- CN202611031373.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-11
- Publication Date
- 2026-09-25
AI Technical Summary
[0047]一、零文字门槛:识字障碍用户只需点击照片即可拨号,完全摆脱文字阅读依赖,操作直观简便。
Smart Images

Figure CN122824841A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of information security technology, specifically relating to a human-computer interaction technology for smart terminals. Background Technology
[0002] With the widespread use of smartphones, mobile communication has become a basic necessity in people's daily lives. However, for special groups such as the elderly, people with sequelae of stroke, people with motor dysfunction, people with literacy difficulties, and people with visual impairments, the existing smartphone interaction methods have serious deficiencies in assistive functions.
[0003] Currently, mainstream smart terminals on the market primarily rely on text-based contact lists, standard touch-screen dialing interfaces, and cloud-based voice assistants (such as Siri and Xiao Ai) to achieve communication functions. Existing technologies have the following shortcomings:
[0004] I. Text-dependent issues. Traditional mobile phone address books display contact names in a text list format, which is incomprehensible to people with literacy or visual impairments, making it impossible for them to make calls independently.
[0005] Second, high touch accuracy is required. The standard dialer's buttons are small, and people with stroke sequelae may have difficulty clicking accurately due to hand tremors or insufficient muscle control, making them prone to accidental touches or inability to complete the operation.
[0006] III. Limitations of Voice Interaction. Existing voice assistants require users to speak the complete and standard name of the contact. For individuals with stroke-related sequelae who cannot clearly express themselves, those with heavy regional accents, and those with speech impairments, the recognition rate is low. Furthermore, most existing voice recognition solutions rely on cloud processing; uploading user voice data to cloud servers poses a risk of privacy breaches.
[0007] Fourth, the system lacks personalized and customizable triggering methods. The existing system does not provide a speed dialing method based on the user's own behavioral characteristics (such as specific tapping patterns or body movements), which fails to meet the special needs of people with physical disabilities.
[0008] Fifth, existing accessibility functions are disconnected from the identity authentication system. Most existing accessibility assistive functions on the market are independent modules that are not integrated with infrastructure such as security chips, electronic signatures, and digital identity authentication, making it impossible to achieve traceability and auditability of communication behavior.
[0009] 6. Inability to proactively send messages when phone calls cannot be connected. People with special needs often do not know how to send text messages or leave messages using instant messaging software when they cannot reach their children or relatives by phone. The existing system lacks an intelligent mechanism to automatically switch from phone calls to message sending when a call is not connected, resulting in the interruption of important communications.
[0010] Therefore, there is an urgent need for a smart terminal-assisted communication method that requires no literacy, no precise touch control, supports personalized customization, can operate offline, can be integrated with the existing digital identity authentication system, and has multiple number attempts and message fallback functions.
[0011] It should be noted that the above description of the technical background is only for the purpose of providing a clear and complete explanation of the technical solutions of the present invention and facilitating understanding by those skilled in the art. It should not be assumed that the above technical solutions are known to those skilled in the art simply because they have been described in the background section of this invention. Summary of the Invention
[0012] The purpose of this invention is to overcome the shortcomings of the prior art and provide a special group of people-assisted communication method and system based on multimodal perception and local trusted execution environment.
[0013] This invention discloses a special-person auxiliary communication system based on multimodal perception and a local trusted execution environment, comprising: a photo dialing unit, used to switch the terminal's address book interface to a photo grid mode, where each contact entry is associated with a photo, and detecting a photo click triggers the system dialing interface to directly dial the corresponding phone number; a local personalized voice dialing library unit, including an encrypted storage area within a security chip, used to store the voiceprint feature vectors of one or more voice samples for each contact, and a local voiceprint comparison engine, wherein the aforementioned voiceprint feature vectors are not uploaded to the cloud; a custom physical trigger unit, used to receive and store the mapping relationship between at least one non-voice physical trigger action set by the user and the contact, wherein the aforementioned physical trigger action includes at least one of screen tap count encoding, device shake mode, touch area coordinates, and external switch signal; and a secondary confirmation unit, used to confirm the caller ID at any trigger. After an event is activated, the system uses voice synthesis to announce the identity information of the contact to be dialed and receives a secondary confirmation signal from the user. The actual dialing is only performed after receiving the confirmation signal. A one-click multi-connection unit stores a priority list of numbers for each contact and automatically attempts to dial multiple numbers sequentially until a connection is established or the list is exhausted. A guaranteed message delivery unit automatically selects or uses a pre-recorded message from a pre-set message template library and sends it to the aforementioned contacts via at least one of SMS, instant messaging applications, or system-level push notifications when all preset phone numbers are unreachable. A communication auditing unit digitally signs the communication information using the electronic signature private key stored in the terminal after each call or message is sent, generating a unique communication signature Seal_ID and storing it. A security management unit interfaces with a five-digit authentication system to verify user identity and manage encryption keys.
[0014] Furthermore, the aforementioned local voiceprint comparison engine employs a voiceprint feature comparison method based on Gaussian Mixture Model-Universal Background Model (GMM-UBM). Specifically, for the i-th speech sample of each contact, the Mel-Frequency Cepstral Coefficients (MFCC) feature vector sequence X is extracted. i ={x1, x2, …, x T The archived model λ is adaptively derived from the general background model using the maximum a posteriori (MAP) probability. i The real-time voice feature vector sequence Y = {y1, y2, ..., y} input by the user. L Match score with the contact archive model (Equation 1)
[0015] In the formula, p is the probability density function, p(y) t |λ i ) indicates that in model λ i The observed eigenvector y t The probability density value.
[0016] When the above matching score exceeds the preset first threshold θ p When a match is found, it is determined that the contact is successfully matched; when multiple contacts have matching scores exceeding θ p When selecting a match, the highest matching score is chosen if the difference between it and the second-highest matching score is greater than a preset second threshold θ. m The contact information is used as the matching result to reduce confusion and misjudgment in voiceprint recognition.
[0017] Furthermore, the aforementioned custom physical trigger unit also includes an action feature template learning module, which performs the following steps: During the configuration phase, it continuously collects three-axis time-series data from the terminal's built-in accelerometer and / or gyroscope when the user performs the same target trigger action multiple times, constructing an original action dataset. A={a1, a2, ..., a N (Equation 2)
[0018] In the formula, a n =(t n acc X (t n ), acc Y (tn ), acc Z (t n ), gyr X (t n ), gyr Y (t n ), gyr Z (t n (Formula 3)
[0019] In the formula, acc represents acceleration. X (t n ) indicates that at t n The acceleration value in the X-axis direction at time acc Y (t n ) indicates that at t n The acceleration value in the Y-axis direction at any given time, acc Z (t n ) indicates that at t n The acceleration value along the Z-axis at any given time; gyr represents the angular velocity of the gyroscope. X (t n ) indicates that at t n The angular velocity value of rotation about the X-axis at any given time, gyr Y (t n ) indicates that at t n The angular velocity value of rotation about the Y-axis at any given time, gyr Z (t n ) indicates that at t n The angular velocity value of rotation around the Z-axis at any given time.
[0020] The above three-axis time series data are segmented using a sliding window. Within each window of length W, the statistical feature vector F = [mean, std, max, min, range, zero-crossing rate, energy] is calculated. Here, mean is the mean of each axis, std is the standard deviation of each axis, max is the maximum value of each axis, min is the minimum value of each axis, range is the range of each axis, zero-crossing rate is the rate at which each axis crosses zero, and energy is the energy value of each axis.
[0021] Dynamic Time Warping (DTW) is used to calculate the current input action sequence Q = {q1, q2, ..., q}. M} and the stored corresponding contact C k Template action sequence P k={p1, p2, ..., p K Optimal path distance between} (Equation 4)
[0022] In the formula, π is the monotonic alignment path and d is the Euclidean distance.
[0023] When the aforementioned dynamic time warping distance is less than the preset third threshold θ d At that time, determine whether the above current input action matches contact C. k The trigger action was successfully matched, triggering a call to contact C. k This allows users to complete dialing operations with natural movements rather than precise touch controls.
[0024] Furthermore, the aforementioned secondary confirmation unit employs a multimodal confirmation signal fusion and determination method based on a time window, specifically including the following steps: After triggering dialing, a time window of length T is initiated. w Within the specified confirmation time window, three confirmation signal sources are simultaneously monitored: touchscreen tap signals, accelerometer shaking signals, and specific voice confirmation words captured by the microphone. Confidence scores for each channel are calculated: tap signal confidence score c. tap The confidence level of the shaking signal, c, is calculated by matching the detected number of taps with the preset number of confirmed taps in the range [0,1]. shake The range ∈[0,1] is obtained by mapping the dynamic time warping distance between the shaking pattern and the preset shaking template to the [0,1] interval, and the confidence level c of the voice confirmation word is obtained. voice ∈[0,1] is obtained from the posterior probability output by the local speech keyword recognition model.
[0025] Calculate the fusion confirmation score C=w1·c tap +w2·c shake +w3·c voice (Equation 5)
[0026] In the formula, w1+w2+w3=1, and each weight is dynamically adjusted according to the user's preset enable status. The weight of a disabled signal source is reset to zero, and the remaining weights are normalized proportionally.
[0027] When the aforementioned fusion confirmation score C exceeds the preset fourth threshold θ c If the second confirmation is successful, the dialing will proceed; otherwise, it will proceed within the time window T. w The dialing will be automatically canceled after the timeout to prevent accidental dialing due to misidentification of single-mode.
[0028] Furthermore, the call failure determination of the aforementioned one-click multi-connection unit adopts a composite judgment method based on call status code and duration as follows: The call status change is monitored through the TelephonyManager interface of the terminal operating system. When the status changes from DIALING to IDLE, the call termination reason code is obtained. If the reason code is any one of CALL_FAIL_UNOBTAINABLE_NUMBER, CALL_FAIL_BUSY, CALL_FAIL_NO_ANSWER, CALL_FAIL_UNREACHABLE, or CALL_FAIL_TIMEOUT, or if the continuous ringing exceeds the preset ringing timeout threshold T... ring If the call never enters OFFHOOK mode, it is determined that the current number is not connected. After determining that the call is not connected, the failure reason code of the current number is recorded, and the system automatically switches to the next number in the priority list to continue trying until a number is connected or the entire list has been tried. Before switching to the next number, a voice synthesis announcement is made saying "Number X not connected, try the next number" to inform the user of the current contact progress.
[0029] Furthermore, the aforementioned message delivery unit includes a message delivery state machine, whose state transition logic is as follows: The initial state is IDLE. When the one-click multi-connection unit reports that all attempts to reach the desired number have failed, the state transitions to MSG_PREPARE. The system selects the message template with the highest matching degree from the preset message template library based on the context. This message template includes replaceable variable placeholders, including the current time, the list of numbers already attempted, and the user's preset nickname. The state transitions to MSG_SENDING. The system attempts to send messages sequentially according to the preset message channel priority list: the first channel is SMS, sent via SmsManager; if SMS sending fails, the system transitions to the second channel, instant messaging application messages, sent via the application's public interface; if the instant messaging application is not installed or sending fails, the system transitions to the third channel, system-level notification push; after any channel successfully sends a message, the state transitions to MSG_SENT, generating a message signature and recording the successful channel type and timestamp; if all channels fail to send a message, the state transitions to MSG_FAILED, the system generates a failure report signature, and notifies the user or their family through a system notification. The aforementioned message template library supports remote updates by family members. Updates take effect after being authenticated by Wushubao, and new templates can only be written to the local template library after being electronically signed.
[0030] Furthermore, the communication signature Seal_ID generated by the aforementioned communication auditing unit adopts a timestamp-based evidence storage method based on blockchain light nodes. Specifically, this includes: after each call or message is sent, the system combines the signature data into a digest message M={user_id, contact_id, timestamp, trigger_type, confirm_type, call_duration / msg_hash, nonce}; using the electronic signature private key to sign the digest message M using the SM2 elliptic curve public key cryptography algorithm to generate a signature value Sig; packaging the hash value H(M) of the digest message M with the signature value Sig, and submitting it to the consortium blockchain network through the blockchain light node client to obtain the block height and transaction hash returned by the blockchain; and associating and storing the block height and transaction hash with the local signature data to form an immutable chain of communication evidence for subsequent traceability and auditing.
[0031] Furthermore, the specific method for the aforementioned security management unit to interface with the five-digit authentication system is as follows: When a user activates the system for the first time, an identity binding request is initiated to the five-digit authentication server by reading the unique device identifier in the terminal security chip and the user's five-digit QR code. After verifying the user's identity, the five-digit authentication server returns the identity credentials, including the user's public key certificate, to the terminal security chip for storage. The aforementioned security chip complies with the national cryptographic level 2 or above of the State Cryptography Administration of the People's Republic of China. All contact photos are stored after symmetric encryption using the AES-256-GCM encryption algorithm. The encryption key is protected by a key derived internally from the security chip, preventing any external process from directly reading the plaintext photo data. Before each dialing attempt, the system verifies whether the current user's identity matches the bound identity stored in the security chip. Only after successful verification can the contact photo and voice sample library be loaded.
[0032] The present invention also discloses a special group-assisted communication method based on multimodal perception and local trusted execution environment, including the following steps.
[0033] Step S1: In response to the first mode switching command, the terminal address book interface is switched to photo grid mode. Each contact entry is associated with a photo imported from the album or taken by the camera. Clicking on the photo will directly trigger the dialing of the corresponding phone number.
[0034] Step S2: Establish a voice sample voiceprint feature vector library for each contact in the terminal's local security chip. Use the Gaussian mixture model-general background model method to compare the user's real-time voice input with the local sample library. If a match is successful, a call is triggered.
[0035] Step S3: Provide a custom trigger event setting interface, receive the mapping relationship between the non-voice physical trigger action and the contact set by the user. The physical trigger action includes at least one of the following: screen tap count, device shake mode, touch area coordinates, and external switch signal. The current input action and the template action are matched and identified through a dynamic time warping algorithm.
[0036] Step S4: After any triggering event is activated, the identity information of the contact to be dialed is broadcast through voice synthesis. Multimodal confirmation signals are collected and fused within a preset time window, and the fused confirmation score is calculated. The actual dialing is only performed after the fused confirmation score exceeds a preset threshold.
[0037] Step S5: In response to dialing trigger, attempt to dial multiple phone numbers sequentially according to a preset priority list. Determine whether each number is connected in real time by monitoring the call status code and ringing duration, until the call is connected or the list is exhausted.
[0038] Step S6: If all preset phone numbers cannot be reached, the message delivery process will be automatically activated. The preset messages will be sent to the above contacts in order of message channel priority via SMS, instant messaging application or system-level push, and the sending status of each channel will be recorded.
[0039] Step S7: After each call or message is sent, use the electronic signature private key stored in the terminal security chip to digitally sign the communication information, generate a unique communication signature Seal_ID, and timestamp it through a blockchain light node and store it in the local secure storage area.
[0040] Furthermore, the screen tap count recognition in step S3 above adopts a time-series pulse detection algorithm based on adaptive threshold, which specifically includes the following steps.
[0041] The touchscreen sensor monitors finger touch and release events, recording a timestamp t for each tap event. k and its pressure value p k This constitutes a knocking event flow. E={e1, e2, ..., e n (Equation 6)
[0042] In the formula, e k =(t k , p k ).
[0043] The event stream is grouped using a sliding time window T_sliding, and the number of taps cnt is counted within the window. When the time interval between two consecutive taps Δt=t k+1 -t k Less than the first interval threshold τ shortJitters determined to be from the same tap are merged; when Δt is greater than the second interval threshold τ, the jitters are merged. long The time is considered the end of a valid tap.
[0044] Introducing a pressure adaptive threshold mechanism: During the configuration phase, a set P of pressure values from n user taps is collected. calib ={p1,p2, ..., p n}, calculate its mean μ p and standard deviation σ p The effective striking condition is the pressure value p. k >μ p -α·σ p , where α is a preset sensitivity coefficient.
[0045] At the end of the detection window, the number of valid taps counted is matched with the number of taps preset for each contact in the configuration table. The contact that matches perfectly and has the smallest deviation in pressure characteristics is selected as the target contact to prevent false triggering caused by non-tapping actions such as light touch or clothing friction.
[0046] The beneficial effects of this invention are as follows.
[0047] 1. Zero literacy barrier: Users with literacy difficulties can make calls simply by clicking on the photo, completely eliminating the need for reading text. The operation is intuitive and simple.
[0048] II. Highly Adaptive Voice Dialing: Supports dialects, fuzzy pronunciations, and incomplete words. Samples are stored locally, recognition does not require an internet connection, protects privacy, and is highly adaptable to special pronunciations.
[0049] III. Multi-mode Customizable Trigger: Offers multiple trigger modes including tapping, shaking, and area touch to accommodate users with varying degrees of physical disability. These modes can also be configured by family members, providing high flexibility. The adaptive pressure calibration threshold dynamically adjusts to address the unstable tapping force of elderly individuals and those with post-stroke sequelae, significantly reducing false triggers.
[0050] IV. Reduced Accidental Operations: Two-level confirmation (voice announcement + secondary confirmation) effectively prevents accidental dialing caused by trembling or unconscious movements. Multimodal fusion judgment makes secondary confirmation more accurate and reliable.
[0051] V. Traceable and auditable: Communication signatures provide complete call credentials, facilitating family monitoring, and blockchain-based evidence storage ensures immutability and complies with data security regulations.
[0052] VI. Seamless integration with existing technology systems: The use of security chips, five-digit data packets, and electronic signatures enhances system security and legal validity.
[0053] VII. Significantly improved contact reliability: The one-click multi-connection mechanism automatically tries multiple numbers to avoid communication interruption due to a single number not working.
[0054] 8. Enhanced passive communication ability: When the phone is not connected, a preset message is automatically sent, enabling special groups to actively transmit contact information even if they do not have the ability to input text, reducing anxiety and the risk of omission. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in an embodiment of the present invention or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. The drawings described below are merely embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0056] Figure 1 This is a schematic diagram of a special population-assisted communication system based on multimodal perception and local trusted execution environment in one embodiment of the present invention.
[0057] Figure 2 This is a flowchart of a special group-assisted communication method based on multimodal perception and local trusted execution environment in one embodiment of the present invention.
[0058] The reference numerals in the above figures are as follows:
[0059] 1. Photo dialing unit; 2. Local personalized voice dialing library unit; 3. Custom physical trigger unit; 4. Secondary confirmation unit; 5. One-click multi-connection unit; 6. Guaranteed message delivery unit; 7. Communication auditing unit; 8. Security management unit; 21. Encrypted storage area; 22. Local voiceprint comparison engine; 81. Five-digit authentication server; 100. Special population auxiliary communication system based on multimodal perception and local trusted execution environment; S1 to S7 are steps. Detailed Implementation
[0060] To better understand this invention, the following embodiments are provided in conjunction with the accompanying drawings. It should be understood that the embodiments of this invention are for illustrative purposes only and not for limiting the invention; the scope of protection of this invention is defined solely by the claims. The embodiments provided are merely preferred embodiments and are not intended to limit the invention in any way. Those skilled in the art can make changes, equivalent substitutions, or modifications based on the content of this invention, resulting in different implementation methods. However, any changes and modifications, or equivalent substitutions, made to the method of this invention without departing from the inventive concept are within the scope of protection of this invention.
[0061] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0062] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, and / or combinations thereof.
[0063] First, please refer to Figure 1 . Figure 1 This is a schematic diagram of a special population-assisted communication system 100 based on multimodal perception and a local trusted execution environment, according to an embodiment of the present invention. Figure 1 As shown, a special population-assisted communication system 100 based on multimodal perception and local trusted execution environment of the present invention includes: a photo dialing unit 1, a local personalized voice dialing library unit 2, a custom physical triggering unit 3, a secondary confirmation unit 4, a one-click multi-connection unit 5, a message delivery guarantee unit 6, a communication auditing unit 7, and a security management unit 8.
[0064] The aforementioned photo dialing unit 1 is used to switch the terminal's address book interface to photo grid mode, where each contact entry is associated with a photo. Clicking on a photo triggers the system dialing interface, directly dialing the corresponding phone number.
[0065] In its implementation, the photo dialing unit 1 expands the terminal's address book database by adding a photo storage path field and a photo ID field to each contact record. It also provides an interactive interface for importing photos from the system album or instantly capturing photos using the terminal's camera. Photo files are stored in encrypted form within an encrypted storage area 21 managed by the terminal's security chip, preventing malicious replacement or unauthorized access. When a user touches the photo display area, the system captures the click event through a touch event listening module, parses the contact ID and phone number associated with the clicked photo, and directly calls the terminal operating system's TelephonyManager interface to initiate a dialing request. This eliminates the need to navigate to the contact details page or a secondary confirmation page, enabling one-click dialing. Simultaneously, the photo dialing unit 1 supports long-pressing a photo to trigger the local speech synthesis engine to announce the contact's name, helping users with mild vision impairment confirm the identity of the clicked contact before use.
[0066] The aforementioned local personalized voice dialing library unit 2 includes an encrypted storage area 21 within a security chip, used to store voiceprint feature vectors of one or more voice samples for each contact, and a local voiceprint comparison engine 22, wherein the voiceprint feature vectors are not uploaded to the cloud.
[0067] In practical implementation, the local personalized voice dialing library unit 2 provides a voice sample recording interface. Users or their family members can access this interface to record multiple voice samples for each contact, including standard Mandarin pronunciation, dialect pronunciation, nicknames, and even incomplete syllables. After preprocessing, each sample extracts a Mel-frequency cepstral coefficient (MFCC) feature vector sequence, which is stored in the encrypted storage area 21 within the secure chip. The local voiceprint comparison engine 22 adopts a Gaussian mixture model-general background model architecture. Using the general background model as a prior, it adaptively derives an archive model from the voice samples of each contact from the general background model using the maximum posterior probability. In an offline terminal environment, it extracts the MFCC features of the user's input voice in real time, calculates the log-likelihood score between the input voice feature sequence and the archive models of each contact, and outputs the contact with the highest score that meets the dual threshold conditions as the matching result. The entire process does not rely on a network connection, ensuring that the user's biometric data is not leaked.
[0068] The aforementioned custom physical trigger unit 3 is used to receive and store the mapping relationship between at least one non-voice physical trigger action set by the user and the contact. The physical trigger action includes at least one of the following: screen tap count encoding, device shake mode, touch area coordinates, and external switch signal.
[0069] In practical implementation, the custom physical trigger unit 3 provides a configuration interface, allowing users or their family members to assign one or more physical trigger methods to each contact. For screen tap count encoding, the system monitors finger contact and release event sequences via touchscreen sensors, counts the number of valid taps within a preset sliding time window, and uses a pressure-adaptive threshold mechanism to determine valid taps. During the configuration phase, the system collects the pressure values of multiple taps from the user and calculates the mean and standard deviation. Only when the tap pressure is greater than the product of the mean, sensitivity coefficient, and standard deviation is it considered a valid tap, thus adapting to individual differences in tapping force among different users. For device shaking mode, the system collects three-axis time-series motion data via accelerometer and gyroscope, extracts statistical feature vectors, and uses a dynamic time warping algorithm to calculate the optimal path distance between the current action sequence and the template action sequence to complete matching and recognition. For touch area coordinates, the system divides the screen into multiple virtual large button areas, each mapped to a different contact. For external switch signals, the system receives trigger signals from external Bluetooth buttons or emergency switches via Bluetooth or wired interfaces, mapping the signal source identifier to the corresponding contact.
[0070] The aforementioned secondary confirmation unit 4 is used to announce the identity information of the contact to be dialed via voice synthesis after any triggering event is activated, and to receive the user's secondary confirmation signal, and to perform the actual dialing only after receiving the confirmation signal.
[0071] In practical implementation, upon detecting any triggering event, the secondary confirmation unit 4 first plays a voice prompt of the contact's name or preset title using the terminal's built-in speech synthesis engine in a local Text-to-Speech (TTS) format. Simultaneously with the voice prompt, the unit initiates a confirmation time window of a preset length, such as three seconds. Within this window, it simultaneously monitors three confirmation signal sources: touchscreen tap signals, accelerometer shaking signals, and specific voice confirmation words captured by the microphone. It calculates the confidence score for each channel, dynamically weights and normalizes the confidence scores according to the user-preset enabled channels, then sums them to obtain the fusion confirmation score. The weights of each disabled signal source are reset to zero, and the remaining weights are re-normalized proportionally. The system only executes the actual dialing if the fusion confirmation score exceeds a preset threshold. If no satisfactory confirmation signal is received within the time window, the system automatically cancels the dialing request.
[0072] The aforementioned one-click multi-connection unit 5 is used to store a priority number list for each contact and automatically attempts to dial multiple phone numbers in sequence when dialing, until the call is connected or the list is exhausted.
[0073] In practice, the one-click multi-connection unit 5 maintains a number sequence list for each contact, with each number entry including a number value, number type, and priority number. Family members can add multiple numbers to each contact and adjust their priority order through the settings interface. When the system determines that a contact needs to be dialed, the one-click multi-connection unit 5 first reads the contact's number list and initiates the call starting with the highest priority number. After each call ends, the unit monitors call status changes through the terminal operating system's TelephonyManager interface. When it detects a change from DIALING to IDLE, it obtains the call termination reason code. If the reason code is any of the preset set of reasons for not being able to connect, or if the ringing continues for more than the preset ringing timeout threshold without entering OFFHOOK state, it determines that the current number is not connected, records the failure reason code, and automatically switches to the next priority number to continue the attempt. Before each switch, the current attempt progress is announced via voice synthesis, informing the user of the contact status.
[0074] The aforementioned message delivery unit 6 is used to automatically select or use a pre-recorded message from the preset message template library when all preset phone numbers cannot be reached, and send it to the contact via at least one of the following methods: SMS, instant messaging application, or system-level push.
[0075] In practical implementation, the message delivery unit 6 maintains a message template library. Family members can pre-enter multiple message templates, each of which can include replaceable variable placeholders, including the current time, a list of numbers already tried, and a user-preset nickname. Internally, the message delivery unit 6 implements a message delivery state machine. Its initial state is IDLE. When the one-click multi-connection unit 5 reports that all attempts have failed, the state transitions to MSG_PREPARE. The system selects a matching message template from the template library based on the context, fills in the variable values, and generates the message to be sent. The state then transitions to MSG_SENDING. This unit attempts to send messages sequentially according to the preset message channel priority order. The first channel is SMS sent via SmsManager. If SMS sending fails, it switches to the second channel: instant messaging application messages sent via the application's public interface. If the instant messaging application is not installed or sending fails, it switches to the third channel: system-level notification push. If a message is successfully sent through any channel, the status changes to MSG_SENT, a message signature is generated, and the channel type and timestamp of the successful message are recorded. If all channels fail to send the message, the status changes to MSG_FAILED, a failure report signature is generated, and the user and / or their family are notified via system notification.
[0076] The aforementioned communication auditing unit 7 is used to digitally sign the communication information using the electronic signature private key stored in the terminal after each call or message is sent, generating a unique communication signature Seal_ID and storing it.
[0077] In practical implementation, after each call ends or message is sent, the communication audit unit 7 automatically collects key information from the communication and combines it into a structured digest. This digest includes at least the user identifier, contact identifier, timestamp, trigger method type, confirmation method type, call duration or message hash value, and a one-time random number to prevent replay attacks. The communication audit unit 7 then uses the electronic signature private key stored in the terminal security chip to sign the digest information using the SM2 elliptic curve public key cryptography algorithm, generating a signature value. Subsequently, the communication audit unit 7 packages the hash value of the digest information with the signature value and submits it to the consortium blockchain network through a blockchain light node client. It obtains the block height and transaction hash returned by the blockchain, associates the block height and transaction hash with the local signature data, and stores them in the local secure storage area 21, forming an immutable and complete chain of communication evidence for subsequent tracing and auditing.
[0078] The aforementioned security management unit 8 is used to interface with the five-digit authentication system, verify user identity, and manage encryption keys. In specific implementation, when a user activates the system for the first time, the security management unit 8 reads the unique device identifier in the terminal's security chip and prompts the user to present their five-digit QR code. By scanning the QR code, the user initiates an identity binding request carrying the device identifier and user identity information to the five-digit authentication server 81. After verifying the user's legitimacy, the five-digit authentication server 81 returns the identity credential, including the user's public key certificate, and writes it into the terminal's security chip for storage. This security chip meets or exceeds the national cryptographic level 2 security standard. Regarding data protection, the security management unit 8 uses the AES-256-GCM symmetric encryption algorithm to encrypt and store all contact photos. The encryption key is protected by a key derived internally from the security chip, preventing any external process from directly reading the plaintext photo data. Before each dialing attempt, the unit verifies whether the current terminal user's identity matches the bound identity stored in the security chip. Only after successful verification is the loading of the contact photo and voice sample library allowed, ensuring that sensitive data remains effectively protected even if the terminal device is lost or obtained by others.
[0079] In one embodiment of the present invention, the local speaker recognition engine 22 first constructs a general background model, which is a Gaussian mixture model trained on a large-scale multi-speaker speech dataset, representing the acoustic feature distribution of a general population. When a user records a voice sample for a contact, the system extracts the Mel-frequency cepstral coefficient feature vector sequence X of that sample. i ={x1, x2, ..., x T}, where T is the number of frames in the feature vector. The system starts with a general background model and uses a maximum a posteriori probability adaptive algorithm to adaptively fine-tune the general background model using the contact's voice feature vector sequence, deriving a unique archive model λ for that contact. i During the adaptive process, the system adjusts only the mean vector while keeping the covariance matrix and weights unchanged, thus obtaining a robust speaker model even with a small number of samples.
[0080] When a user issues a voice command during actual use, the system collects the user's voice in real time and extracts the Mel-frequency cepstral coefficient feature vector sequence Y={y1, y2, ..., y L}, where L is the number of feature vector frames for the current input speech. For each contact archive model λ in the repository i The system calculates the average log-likelihood score of the input speech feature sequence Y under this model. (Equation 1)
[0081] In the formula, p is the probability density function, p(y) t |λ i ) indicates that in model λ i The observed eigenvector y t The probability density value.
[0082] The average log-likelihood score reflects the degree of matching between the current input speech features and the acoustic model of the i-th contact. During the matching decision phase, the system employs a dual-threshold decision strategy. First, a first threshold θ is set. p As a basic matching threshold, when a contact's matching score exceeds θ p When a match is found, the contact is added to the candidate set. If there is only one contact in the candidate set, that contact is directly output as the match result. If there are multiple contacts in the candidate set, the system further calculates the difference between the highest and second-highest matching scores in the candidate set. When this difference is greater than a preset second threshold θ, the match is considered complete. m When the highest score is selected, the contact corresponding to that score is chosen as the final matching result. This dual-threshold strategy is specifically designed for communication scenarios involving a limited number of contacts but potentially similar acoustic features, effectively reducing the risk of mis-dialing contacts due to voiceprint feature confusion.
[0083] It is worth noting that, in terms of storage protection, all the archived model parameters of all contacts, namely the mean vector, covariance matrix and weight coefficient of each Gaussian component, as well as the user's real-time voice feature vector, are stored only in the encrypted storage area 21 managed by the security chip. The entire extraction, comparison and decision-making process is completed offline on the terminal and no data is transmitted to any cloud server.
[0084] In one embodiment of the present invention, the action feature template learning module of the custom physical trigger unit 3 provides an action input guidance interface for the user during the configuration phase. When the user is about to set a trigger action for a certain contact, the special population assistance communication system 100 prompts the user to perform the same target trigger action multiple times consecutively, such as performing the action of "quickly shaking the phone twice" five times consecutively. During the user's execution of the action, the accelerometer built into the terminal continuously collects three-axis acceleration data at a preset sampling frequency, such as 50Hz to 100Hz, and the gyroscope collects three-axis angular velocity data at the same time. The special population assistance communication system 100 combines the raw time series data collected by the sensors into a structured action dataset A={a1, a2, ..., a N},
[0085] In the formula, a n =(t n acc X (t n), acc Y (t n ), acc Z (t n ), gyr X (t n ), gyr Y (t n ), gyr Z (t n (Formula 3)
[0086] In the formula, acc represents acceleration. X (t n ) indicates that at t n The acceleration value in the X-axis direction at time acc Y (t n ) indicates that at t n The acceleration value in the Y-axis direction at any given time, acc Z (t n ) indicates that at t n The acceleration value along the Z-axis at any given time; gyr represents the angular velocity of the gyroscope. X (t n ) indicates that at t n The angular velocity value of rotation about the X-axis at any given time, gyr Y (t n ) indicates that at t n The angular velocity value of rotation about the Y-axis at any given time, gyr Z (t n ) indicates that at t n The angular velocity value of rotation around the Z-axis at any given time.
[0087] Where each data point a n Including timestamp t n In addition, the corresponding three-axis acceleration and three-axis angular velocity values at each moment, totaling nine dimensions of motion data.
[0088] After completing the initial data acquisition, the special population assisted communication system 100 performs sliding window segmentation on the triaxial time series data of each action sample. The length of each window, W, is set to 1 second, for example. Adjacent windows can be partially overlapped to improve the continuity of feature extraction. Within each window length W, the system calculates a multidimensional statistical feature vector F=[mean,std,max,min,range,zero-crossing rate,energy]. This feature vector includes the mean, standard deviation, maximum value, minimum value, range, zero-crossing rate, and energy of each axis. These statistical features comprehensively characterize the amplitude, fluctuation, and frequency characteristics of the action in the time domain.
[0089] When a user performs a physically triggered action in actual use, the special needs assistive communication system 100 processes the currently input action sequence Q={q1, q2, ..., q} according to the same sampling and feature extraction process. M}. For the repository corresponding to contact C k Template action sequence P k ={p1, p2, ..., p K The system uses a dynamic time warping algorithm to calculate Q and P. k The optimal path distance between the two points. The Dynamic Time Warping algorithm constructs an M·K distance matrix and searches for a monotonically aligned path π from the starting point to the ending point within the matrix, such that the sum of the accumulated distances along this path is minimized. (Equation 4)
[0090] In the formula, d is the Euclidean distance metric.
[0091] The special population assisted communication system 100 compares the calculated dynamic time-warped distance with a preset third threshold θ. d Compare the distances. When the distance is less than θ... d At that time, determine whether the current input action matches the contact C. k The trigger action template matched successfully, triggering a call to contact C. k The phone number. This dynamic time warping matching method is naturally tolerant of stretching and compression of movements on the time axis, and can effectively adapt to the problem of uneven speed and rhythm changes when performing the same movement each time for users with limb disabilities such as stroke sequelae and Parkinson's patients. This allows users to complete the dialing operation with natural movements rather than precise touch.
[0092] In one embodiment of the present invention, after the special population assisted communication system 100 detects any dialing trigger event, the secondary confirmation unit 4 immediately initiates a process of length T. w The confirmation time window is set, for example, with a window length of 3 seconds. Within this time window, the assistive communication system 100 for special needs simultaneously activates three confirmation signal monitoring channels: a touchscreen tap signal monitoring channel, an accelerometer shaking signal monitoring channel, and a microphone voice confirmation word signal monitoring channel. These three channels operate independently and in parallel, without interference. For the touchscreen tap signal channel, the assistive communication system 100 monitors finger contact and release events on the screen using the touchscreen sensor, counts the number of valid taps within the time window, and compares the detected tap count with the user-preset confirmation tap count to calculate the tap signal confidence level c. tap Its value ranges from 0 to 1, and is a continuous value. The confidence level approaches 1 when the detected number of taps exactly matches the preset number; as the deviation in the number of taps increases, the confidence level decreases accordingly. Furthermore, for the accelerometer shaking signal channel, the system uses a dynamic time warping method to compare the shaking motion sequence collected within the current time window with the user-preset confirmed shaking template, calculates the dynamic time warping distance, and converts this distance into a shaking signal confidence level c in the 0-1 interval through a nonlinear mapping function. shake The smaller the distance, the higher the confidence level.
[0093] It is worth noting that for the microphone voice confirmation signal channel, the special needs assistive communication system 100 runs a lightweight voice keyword recognition model locally. This voice keyword recognition model adopts a wake-word detection architecture based on a deep neural network and processes the audio stream input from the microphone in real time locally on the terminal. When the model detects user-preset confirmation keywords such as "okay," "call," "confirm," or specific dialect words, it outputs the posterior probability of that keyword as the voice confirmation word confidence level c. voice After obtaining the confidence scores of the three channels, the special needs assistive communication system 100 determines the activation status of each channel according to the user's preset configuration.
[0094] For example, the voice channel is disabled for users with speech impairments who cannot speak; the tapping channel may be disabled for users with limited hand movement. The system sets the weights w1, w2, or w3 corresponding to each disabled channel to zero, and the weights of the remaining undisabled channels are renormalized proportionally to ensure that the sum of the weights of all enabled channels is always equal to 1.
[0095] The system then calculates and confirms the score. C=w1·c tap +w2·c shake +w3·c voice (Equation 5)
[0096] The special population assistive communication system 100 compares the calculated fusion confirmation score C with a preset fourth threshold θ c Comparison, when C exceeds θ c If the second confirmation is successful, the special population auxiliary communication system 100 will execute the actual dialing operation. If within time window T... w If the timeout confirmation score still does not reach the threshold, this invention automatically cancels the current dialing request and plays a cancellation prompt voice.
[0097] This multimodal fusion mechanism enables the invention to accurately identify the user's confirmation intent through comprehensive judgment of other modalities, even if a certain modality cannot provide a reliable confirmation signal due to the user's physiological limitations. This effectively solves the problem of insufficient reliability of single-modal confirmation methods in applications involving special populations.
[0098] In one embodiment of the present invention, after each call is initiated, the one-click multi-connection unit 5 registers a call status listener through the TelephonyManager interface of the terminal operating system to continuously track the status changes of the current call. When the special population assistance communication system 100 dials a number, the call status first enters the DIALING state to indicate that it is dialing. Subsequently, if the other party answers, the status switches to OFFHOOK to indicate that the call is connected. If the other party does not answer, the status may directly switch from DIALING back to IDLE to indicate that the dialing has ended. When the special population assistance communication system 100 detects that the call status has changed from DIALING to IDLE, it triggers the call failure determination process. The system first calls the TelephonyManager's call termination reason code acquisition interface to obtain an integer value reason code, which is generated by the terminal's underlying telephone protocol stack according to the specific reason for the call failure. The special needs assistive communication system 100 compares the reason code with a preset set of unreachable reason codes. This set includes: CALL_FAIL_UNOBTAINABLE_NUMBER (number unreachable), CALL_FAIL_BUSY (the other party is busy), CALL_FAIL_NO_ANSWER (no answer from the other party), CALL_FAIL_UNREACHABLE (the other party is not in the service area), and CALL_FAIL_TIMEOUT (call timeout). When the obtained reason code belongs to any of the above sets, the special needs assistive communication system 100 directly determines that the current number is not connected.
[0099] In some cases where the system may not return a clear reason code, it also employs an auxiliary determination method based on ringing duration. After dialing, the system starts a timer to continuously monitor whether the call status enters OFFHOOK mode. If continuous ringing exceeds a preset ringing timeout threshold T... ringFor example, if the timer is set to 30 seconds or 45 seconds and no OFFHOOK status is detected, the system will also determine that the current number is not connected and forcibly end the call attempt.
[0100] Once a call is determined to be unanswered, the system records the failure reason code for subsequent auditing and log review. It then automatically retrieves the next priority number for that contact from the one-click multi-connection unit 5 and switches to that number to initiate a new call. Before switching to the next number, the system uses the terminal's speech synthesis engine to issue a voice prompt, such as "The first number is unanswered; try dialing the second number," informing the user that a switching of contact channels is underway and preventing anxiety from prolonged waiting. This composite determination method fully covers various unanswered situations that may occur in real-world communication scenarios, ensuring that the one-click multi-connection unit 5 can accurately identify failure states and promptly advance the contact process.
[0101] In one embodiment of the present invention, the message delivery unit 6 internally implements a finite state machine to manage the complete lifecycle of message delivery. This state machine includes five states: IDLE, MSG_PREPARE, MSG_SENDING, MSG_SENT, and MSG_FAILED. The state machine initially operates in the IDLE state. When the one-click multi-connection unit 5 reports that all attempts to connect to its priority list have failed, the state machine transitions from the IDLE state to the MSG_PREPARE state.
[0102] In the MSG_PREPARE state, the system selects the message template with the highest matching degree from the preset message template library based on the current context information. The message template library stores multiple preset text templates, each including text content and applicable condition tags. The system selects the template with the highest tag matching degree based on context information such as the currently tried number list and the current time. The selected template may include replaceable variable placeholders, such as "current time" placeholders, "tryed number list" placeholders, and "user preset nickname" placeholders. The system replaces these placeholders with actual values to generate the complete message text to be sent. The "user preset nickname" placeholder in the message template refers to the system allowing family members to preset a title for each user in the configuration interface, such as a kinship term or nickname. This placeholder is replaced with the corresponding title text during template rendering.
[0103] After the message text is generated, the state machine transitions to the MSG_SENDING state. In this state, the system attempts to send the message sequentially according to a preset message channel priority list. The first channel is SMS, where the system sends an SMS to the contact's primary mobile number via the SmsManager interface provided by the Android operating system or the MFMessageComposeViewController interface provided by the iOS operating system. If SMS sending fails, the system switches to the second channel, sending an in-app text message via the public content interface or Intent mechanism provided by an installed instant messaging application such as WeChat or DingTalk. If the instant messaging application is not installed or sending fails, the system switches to the third channel, sending a push notification message via an operating system-level notification push service such as Firebase Cloud Messaging or Huawei Push Kit.
[0104] After a message is successfully sent through any of the aforementioned channels, the state machine transitions to the MSG_SENT state. The Assistive Communication System for Special Populations 100 generates a message signature including the channel type and timestamp of successful transmission, recording complete information about this message transmission. If all preset channels fail to send the message, the state machine transitions to the MSG_FAILED state. The Assistive Communication System for Special Populations 100 generates a failure report signature and displays a failure notification to the user and / or their family members through the terminal system notification bar.
[0105] It is worth noting that regarding the update management of the message template library, the Special Population Assistive Communication System 100 supports family members updating the message template library remotely. When a family member initiates a template update request, the request is first verified by the five-digit authentication system to confirm the operator's identity as the user's legal guardian. Only after confirming that the requester is the legal guardian can the requester submit the new template content. Before the new template is written to the local template library on the terminal, it must be digitally signed using the submitter's electronic signature private key. The system only writes the new template to the template library after verifying the signature's validity at the receiving end, thereby ensuring the authenticity of the template content's source and its non-repudiation.
[0106] In one embodiment of the present invention, the communication auditing unit 7 automatically triggers the evidence storage process after each call or message is sent. The special population assisted communication system 100 first collects key information of the current communication event and combines it into structured summary information M={user_id, contact_id, timestamp, trigger_type, confirm_type, call_duration / msg_hash, nonce}. Wherein, user_id is the unique identifier of the current user in the system, contact_id is the unique identifier of the contact person of the communication partner, timestamp is the timestamp of the communication, trigger_type is the triggering method of this communication, such as photo click, voice matching, or tapping action, confirm_type is the secondary confirmation method, such as tap confirmation or voice confirmation, the call_duration / msg_hash field is filled with the hash value of the call duration or message content according to the communication type, and nonce is a randomly generated one-time random number used to prevent replay attacks. The special population assisted communication system 100 encodes the aforementioned digest information M according to a predetermined format, then calls the electronic signature private key stored in the terminal security chip and uses the SM2 elliptic curve public key cryptography algorithm issued by the State Cryptography Administration of the People's Republic of China to digitally sign the digest information M, generating a signature value Sig. The signature value output by the SM2 asymmetric encryption signature algorithm includes the complete mathematical correlation of the digest information; any tampering with the digest information will result in signature verification failure.
[0107] Subsequently, the special population assisted communication system 100 performs a hash operation on the digest information M to obtain a hash value H(M). H(M) and the signature value Sig are then packaged into a single notarized transaction according to a predefined communication protocol format. The special population assisted communication system 100 submits this notarized transaction to the consortium blockchain network through a blockchain light node client built into the terminal. The light node client only stores the block header information and not the complete block data, thus enabling it to operate normally with the limited storage resources of the terminal. After the consensus nodes in the consortium blockchain network verify the validity of the signature in the notarized transaction, they package the transaction into a new block and return the block height and corresponding transaction hash of the notarized transaction.
[0108] The Special Population Assisted Communication System 100 associates and stores the received block height and transaction hash with locally generated signature data. In subsequent tracing, users or their families can retrieve the corresponding evidence record using the block height and transaction hash in a blockchain explorer or consortium blockchain query interface to verify the integrity and temporal authenticity of the communication record. This blockchain evidence storage mechanism ensures the immutability and verifiability of the communication evidence chain, providing legally valid electronic evidence for potential disputes. Simultaneously, the nonce field in the digest information ensures that even two communication records with identical content will generate different hash values for their digest information, thus avoiding evidence confusion caused by duplicate hash values. Each evidence storage operation generates a new random nonce value, mathematically guaranteeing the uniqueness of each communication signature.
[0109] In one embodiment of the present invention, the security management unit 8 executes an initial identity binding process when the user activates the system for the first time. The special population assistance communication system 100 reads the unique device identifier in the built-in security chip through the security interface provided by the terminal operating system. This identifier is written into the chip during production and cannot be tampered with. At the same time, the system prompts the user to present their digital identity QR code. The digital identity QR code is a standardized QR code credential that includes personal or organizational digital identity information, where the individual identity QR code represents personal identity, the family identity QR code represents family identity, the enterprise identity QR code represents enterprise identity, the social identity QR code represents social organization identity, and the government identity QR code represents government agency identity. After the user scans the QR code with the terminal camera, the special population assistance communication system 100 combines the device identifier with the user identity information read from the QR code to construct an identity binding request, and sends the request to the digital identity authentication server 81 via the secure HTTPS protocol.
[0110] After receiving the binding request, the five-digit authentication server 81 verifies the validity of the user's identity information in the request. Once the user's identity is confirmed to be legitimate, it generates a digital identity credential for the user, including the user's public key certificate. This digital identity credential uses the X.509 standard format and includes the user's identity information, public key, and the authentication server's digital signature. The five-digit authentication server 81 returns the identity credential to the terminal through a secure channel. The special needs assistance communication system 100 receives the credential and writes it into the dedicated storage area 21 of the security chip. This security chip conforms to the national cryptographic level 2 or higher security standards promulgated by the State Cryptography Administration of the People's Republic of China and possesses physical tamper-proof and logical attack-proof capabilities.
[0111] It is worth noting that, regarding data encryption protection, the Special Needs Assistive Communication System 100 uses the AES-256-GCM symmetric encryption algorithm to encrypt and store each contact photo. AES-256-GCM is a widely validated authentication encryption mode that generates an authentication tag while encrypting data, ensuring both data confidentiality and integrity. The symmetric key used for encryption is not stored directly, but is derived from a root key within the security chip through a key derivation function, providing key protection.
[0112] Specifically, the system first generates a data encryption key for encrypting photos. This key is then encrypted by a key encryption key and stored in secure storage area 21. When decryption is needed, the security chip performs the decryption operation internally. The plaintext value of the data encryption key never leaves the security boundary of the security chip, and no external process can directly obtain this key. Furthermore, regarding authentication, the special needs assistive communication system 100 performs identity verification before each dialing attempt. The system captures the current user's real-time facial image using the terminal's front-facing camera, compares the feature vector of the facial image with the biometric template of bound users stored in the security chip, or requires the user to enter a PIN code and pass verification through the security chip to confirm that the current user is the legitimate user corresponding to the identity credential in the security chip. Only after successful verification is the system allowed to load sensitive data such as contact photos and voice sample libraries. If verification fails, access is denied and an abnormal attempt log is recorded. This mechanism ensures that even if the terminal device is lost or obtained by others, contact data is still protected by both the security chip and authentication.
[0113] Please refer to Figure 2 . Figure 2 This is a flowchart of a special group-assisted communication method based on multimodal perception and a local trusted execution environment, according to an embodiment of the present invention. Figure 2 As shown, the auxiliary communication method of the above-mentioned special population assistance system includes the following steps.
[0114] Step S1: In response to the first mode switching command, the terminal address book interface is switched to photo grid mode. Each contact entry is associated with a photo imported from the album or taken by the camera. Clicking on the photo will directly trigger the dialing of the corresponding phone number.
[0115] Step S2: Establish a voice sample voiceprint feature vector library for each contact in the terminal's local security chip. Use the Gaussian mixture model-general background model method to compare the user's real-time voice input with the local sample library. If a match is successful, a call is triggered.
[0116] Step S3: Provide a custom trigger event setting interface, receive the mapping relationship between the non-voice physical trigger action and the contact set by the user. The physical trigger action includes at least one of the following: screen tap count, device shake mode, touch area coordinates, and external switch signal. The current input action and the template action are matched and identified through a dynamic time warping algorithm.
[0117] Step S4: After any triggering event is activated, the identity information of the contact to be dialed is broadcast through voice synthesis. Multimodal confirmation signals are collected and fused within a preset time window, and the fused confirmation score is calculated. The actual dialing is only performed after the fused confirmation score exceeds a preset threshold.
[0118] Step S5: In response to dialing trigger, attempt to dial multiple phone numbers sequentially according to a preset priority list. Determine whether each number is connected in real time by monitoring the call status code and ringing duration, until the call is connected or the list is exhausted.
[0119] Step S6: If all preset phone numbers cannot be reached, the message delivery process will be automatically activated. The preset messages will be sent to the above contacts in order of message channel priority via SMS, instant messaging application or system-level push, and the sending status of each channel will be recorded.
[0120] Step S7: After each call or message is sent, use the electronic signature private key stored in the terminal security chip to digitally sign the communication information, generate a unique communication signature Seal_ID, and timestamp it through a blockchain light node and store it in the local secure storage area.
[0121] In one embodiment of the present invention, Grandma Liu, 82 years old, is illiterate, lives alone, and has relatively good eyesight but her fingers are immobile due to arthritis. She needs to frequently contact her son, daughter, and neighbor Aunt Wang. Grandma Liu cannot read the text in her address book, nor can she accurately click small buttons, but she can recognize photos and perform relatively large-scale screen tapping operations.
[0122] The system configuration (one-time preset) is as follows. Family members pre-configure the following: In the photo dialing unit 1, import a clear, front-facing photo of the "son"; in the custom physical trigger unit 3, set a single tap for confirmation (for secondary confirmation in step S4); in the one-click multi-connection unit 5, preset three numbers for the "son": first priority son's mobile phone 138****0001, second priority son's landline 010****5678, third priority daughter-in-law's mobile phone 139****1111; in the guaranteed message delivery unit 6, preset the message template: "I am Grandma Liu, I tried calling you but couldn't get through. Please call back after seeing this"; the security management unit 8 has completed identity authentication, and the photo and key are stored in the security chip.
[0123] The dialing procedure is as follows.
[0124] Step S1: Grandma Liu wants to call her son. She sees a large photo of him on her phone screen and taps the photo area with her finger. The photo dialing unit 1 captures the tap event through the touch event listening module and parses the contact ID associated with the photo as "son" and the corresponding phone number 138****0001. The photo dialing unit 1 passes the trigger event to the secondary confirmation unit 4, and the system simultaneously announces through voice synthesis, "You will dial your son's number."
[0125] Step S2: In this embodiment, Grandma Liu used the photo dialing method and did not use the voice dialing function, so step S2 was not triggered.
[0126] Step S3: In this embodiment, the tapping mapping configuration in step S3 is used for secondary confirmation in step S4. When Grandma Liu taps the screen, the custom physical trigger unit 3 uses an adaptive threshold timing pulse detection algorithm to collect the timestamp and pressure value of the tapping event. After merging the dual time interval threshold jitter and validly determining the pressure adaptive threshold, one valid tap is identified.
[0127] Step S4: After receiving the trigger event from the photo dialing unit 1, the secondary confirmation unit 4 initiates the process for length T. w A 3-second confirmation window is provided, and a voice synthesis announcement is made: "We will be calling your son. Please tap the screen to confirm." After hearing the announcement, Grandma Liu taps the screen anywhere on the screen. The secondary confirmation unit 4 receives the tap signal from the custom physical trigger unit 3 within the confirmation window, counting one valid tap, perfectly matching the preset number of confirmation taps. The tap signal confidence level is c. tap =0.95. Since both the shake confirmation channel and the voice confirmation channel are disabled, the fusion confirmation score C = 0.95. Preset fourth threshold θ c =0.70, C=0.95>θ c The secondary confirmation unit 4 determines that the confirmation is successful and transmits the confirmation signal and contact information to the one-touch multi-connection unit 5. If Grandma Liu does not tap the screen within 3 seconds, the system will automatically cancel the dialing request after the time window expires.
[0128] Step S5: After receiving the dialing command, the one-touch multi-connection unit 5 reads the priority number list of the "son's" contacts: Number 1 is 138****0001 (priority 1), Number 2 is 010****5678 (priority 2), and Number 3 is 139****1111 (priority 3). The one-touch multi-connection unit 5 first dials number 1 through the TelephonyManager interface of the terminal operating system. Number 1 rings continuously for about 20 seconds before its status changes from DIALING to IDLE, never entering OFFHOOK mode. The one-touch multi-connection unit 5 obtains the call termination reason code, and the return value is CALL_FAIL_NO_ANSWER, indicating that number 1 was not connected. The system announces via voice synthesis, "The son's mobile phone is not answering; we are trying to dial the landline." The one-touch multi-connection unit 5 automatically switches to number 2 and initiates a call to 010****5678. Number 2 rings for about 5 seconds before being successfully connected, and the call status enters OFFHOOK mode. Grandma Liu and her daughter-in-law talk for about 3 minutes before hanging up. After the call ends, the one-click multi-connection unit 5 transmits the communication result (connected number 010****5678, call duration 3 minutes) to the communication audit unit 7.
[0129] Step S6: Since the second number was successfully connected in step S5, this step is not triggered. If all three number attempts fail in step S5, the message delivery unit 6 will be activated. Its state machine will transition from IDLE to MSG_PREPARE to select a message template, and then to MSG_SENDING to attempt to send the message via SMS, instant messaging applications, system-level push notifications, etc. After successful sending, it will transition to MSG_SENT to generate a message signature.
[0130] In step S7, after the call ends, the communication audit unit 7 automatically collects key information from the communication and combines it into a digest message M={user_id: "U1005", contact_id: "son", timestamp: "2026-07-09T09:30:00Z", trigger_type: "photo_tap", confirm_type: "tap", call_duration: "180s", nonce: "f8d3a9e2"}. The communication audit unit 7 calls the electronic signature private key stored in the security management unit 8 and uses the SM2 elliptic curve public key cryptography algorithm to sign the digest message M, generating a signature value Sig. Then, the hash value H(M) of the digest message M is packaged with the signature value Sig and submitted to the consortium blockchain network through a blockchain light node client to obtain the block height and transaction hash, which is then stored in association with the local signature data.
[0131] It is worth noting that, in one embodiment of the present invention, the screen tap count recognition in step S3 above adopts a time-series pulse detection algorithm based on adaptive threshold, which specifically includes the following steps.
[0132] The touchscreen sensor monitors finger touch and release events, recording a timestamp t for each tap event. k and its pressure value p k This constitutes a knocking event flow. E={e1, e2, ..., e n (Equation 6)
[0133] In the formula, e k =(t k , p k ).
[0134] Using a sliding time window T sliding Group the event stream and count the number of taps cnt within the window. When the time interval between two consecutive taps is Δt=t... k+1 -t k Less than the first interval threshold τ short Jitters determined to be from the same tap are merged; when Δt is greater than the second interval threshold τ, the jitters are merged. long A valid tap is considered complete when the tap is timed out. A pressure-adaptive threshold mechanism is introduced: during the configuration phase, a set P of pressure values from n taps by the user is collected. calib ={p1, p2, ..., p n}, calculate its mean μ p and standard deviation σ p The effective striking condition is the pressure value p. k >μ p -α·σ p Where α is the preset sensitivity coefficient. At the end of the detection window, the number of valid taps is matched with the preset number of taps for each contact in the configuration table, and the contact with the perfect match and the smallest deviation in pressure characteristics is selected as the target contact to prevent false triggering caused by non-tapping actions such as light touch or clothing friction.
[0135] For example, a touchscreen sensor continuously collects touch data at a system sampling rate, such as 100Hz. It records a touch start event whenever a finger is detected touching the screen and a touch end event when the finger is detected leaving the screen. The system pairs a touch start event with the immediately following touch end event to form a complete tap event, where the timestamp t... k Take the time when the touch end event occurs and the pressure value p. kThe maximum or average pressure value reported by the sensor during the touch process is retrieved. The events in this event stream are arranged in chronological order, forming the basic input data for subsequent timing pulse detection and processing.
[0136] The aforementioned custom physical trigger unit 3 uses a sliding time window T. sliding Group the tapping event stream and count the number of valid taps within a window.
[0137] In practice, the system maintains a fixed-length sliding time window, for example, 800 milliseconds. This window continuously slides forward as new tap events arrive. Whenever a new tap event is recorded, the system includes it in the current time window and discards historical events that exceed the window's initial boundary, ensuring that the window always includes the most recent T. sliding All tapping events within the time period. The system performs time-series analysis on the events within this window, counting the number of valid taps as the basis for subsequent matching. The aforementioned custom physical trigger unit 3 analyzes the time intervals of consecutive tapping events within the window pairwise. When the time interval between two consecutive taps Δt = t k+1 -t k Less than the first interval threshold τ short Vibrations that are judged to be from the same tap are merged.
[0138] In practical implementation, the system presets a first interval threshold τ. short For example, 100 milliseconds represents the upper limit of the time interval that is impossible in a normal tapping action. When the time interval between two consecutive taps is less than this threshold, it indicates that these two events are not two independent taps consciously performed by the user, but rather redundant events generated during the same intended tapping process due to factors such as finger tremor, screen bounce, or sensor noise. The system merges such events into a single valid tap, using the event with the highest pressure value in the group as the representative event, and retaining its timestamp and pressure value as attributes of this valid tap.
[0139] When the time interval Δt between two consecutive taps is greater than the second interval threshold τ long At that time, the custom physical trigger unit 3 determines that a valid tap has ended.
[0140] In practical implementation, the system presets a second interval threshold τ. long For example, 400 milliseconds represents a reasonable time interval between the completion of one valid tap and the start of the next tap in a normal tapping action. When the time interval between two consecutive tapping events is greater than this threshold, it indicates that the previous tap has completely ended, and the next tap is the start of an independent new tap. Based on this, the system assigns the two events to different valid taps, thus correctly counting the number of independent taps. For Δt between τ... shortand τ long In the case of the two events, the system determined that they were two independent taps, but the time interval was normal, and recorded them as two valid taps.
[0141] The aforementioned custom physical trigger unit 3 introduces a pressure adaptive threshold mechanism, which collects a set of pressure values P from multiple taps by the user during the configuration phase. calib ={p1, p2, ..., p n}, calculate its mean μ p and standard deviation σ p The effective striking condition is the pressure value p. k >μ p -α·σ p , where α is a preset sensitivity coefficient.
[0142] In practical implementation, the system guides the user to perform multiple tapping actions during the configuration phase, such as performing 10 standard taps consecutively, and collecting the pressure value of each tap to form a calibration set P. calib The system calculates the arithmetic mean μ of the calibration set. p and standard deviation σ p The system calculates the effective tapping pressure threshold based on a preset sensitivity coefficient α. The sensitivity coefficient α typically ranges from 0.3 to 0.8, preferably 0.5, and controls the leniency of the effective tapping judgment: the larger the α value, the lower the pressure threshold and the more lenient the judgment, suitable for users with lighter tapping force; the smaller the α value, the higher the pressure threshold and the stricter the judgment, suitable for users with stable tapping force who need to strictly prevent accidental triggering.
[0143] In practical use, if the pressure value of any touch event is lower than the adaptive threshold p, threshold =μ p -α·σ p If a touch is not detected, it will be considered an invalid touch and will not be counted in the number of taps, thus effectively eliminating interference signals caused by clothing rubbing, accidental touches by the palm, or unintentional light touches.
[0144] At the end of the detection window, the custom physical trigger unit 3 matches the number of valid taps counted with the number of taps preset for each contact in the configuration table, and selects the contact that matches perfectly and has the smallest deviation in pressure characteristics as the target contact.
[0145] In practical implementation, the system operates within a sliding time window T. sliding At the end, the system obtains the number of valid taps N counted within the window. The system reads the mapping relationship between each contact and its preset tap count stored in the custom physical trigger unit 3 configuration table, and filters out all contacts whose preset tap count is exactly the same as N as the candidate set.
[0146] If there is only one contact in the candidate set, the system directly selects that contact as the target contact. If there are multiple contacts in the candidate set, the system further calculates the average value p of each effective tapping pressure value within the current detection window. mean The pressure deviation is compared with the average pressure calibration value when each contact is configured in the candidate set, and the pressure deviation |p is selected. mean -μ p(k) The smallest contact is taken as the target contact, where μ p(k) This is the average pressure calibration value for the k-th contact during the configuration phase. This dual-selection strategy ensures that the system can still effectively distinguish between multiple contacts by the pressure feature dimension when multiple contacts are preset with the same number of taps, further improving the accuracy and robustness of the identification.
[0147] In the aforementioned custom physical trigger unit 3, the complete execution flow of the screen tap count recognition algorithm also includes a wake-up screen state check.
[0148] In practice, the system only processes tap events when the terminal screen is awake. When the screen is off or locked, all touch events are ignored to prevent accidental triggering due to bumps when the terminal is in a pocket or bag. The system also requires that the first valid tap within the detection window must occur within a preset time window after the screen is awakened. Taps that exceed this time are not recognized as valid dialing triggers and require the screen to be awakened again via a physical button or specific gesture before they can be tried again.
[0149] In addition, the custom physical trigger unit 3 analyzes the duration of touch events to distinguish between tapping and swiping operations.
[0150] In practice, the system records the duration of each touch event, i.e., the time difference between the start and end of the touch. When this duration exceeds a preset touch duration threshold, such as 500 milliseconds, the system determines that the event is a swipe or long press rather than a tap, excludes the event from the tap event stream, and does not participate in subsequent tap counts. Through this mechanism, the system can effectively prevent users from mistakenly identifying dragging or swiping on the screen as tapping operations.
[0151] The core of this invention lies in providing a smart assisted communication method and system for special groups that requires no literacy, no precise touch control, supports personalized multimodal triggering methods, has automatic multi-number tracing and message fallback functions, and protects privacy throughout the process while being traceable and auditable.
[0152] The embodiments of the present invention described above can be implemented in various hardware, software codes, or combinations thereof. For example, embodiments of the present invention can also be program code executing the above methods in a Digital Signal Processor (DSP). The present invention can also relate to various functions executed by a computer processor, digital signal processor, microprocessor, or Field Programmable Gate Array (FPGA). The processor described above can be configured to perform specific tasks according to the present invention, which are accomplished by executing machine-readable software code or firmware code defining the specific methods disclosed in the present invention. The software code or firmware code can be developed into different programming languages and different formats or forms. The software code can also be compiled for different target platforms. However, the different code styles, types, and languages of the software code performing tasks according to the present invention and other types of configuration code do not depart from the spirit and scope of the present invention.
Claims
1. A special-person assisted communication system based on multimodal perception and a local trusted execution environment, comprising: The photo dialing unit is used to switch the terminal's address book interface to photo grid mode, where each contact entry is associated with a photo. Clicking on a photo triggers the system dialing interface to directly dial the corresponding phone number. The local personalized voice dialing library unit includes an encrypted storage area within the security chip and a local voiceprint comparison engine. The encrypted storage area is used to store the voiceprint feature vectors of one or more voice samples for each contact, and the voiceprint feature vectors are not uploaded to the cloud. A custom physical trigger unit is used to receive and store the mapping relationship between at least one non-voice physical trigger action set by the user and the contact. The physical trigger action includes at least one of the following: screen tap count encoding, device shake mode, touch area coordinates, and external switch signal. The secondary confirmation unit is used to announce the identity information of the contact to be dialed via voice synthesis after any triggering event is activated, and to receive the user's secondary confirmation signal. It will only perform the actual dialing after receiving the confirmation signal. The one-click multi-connection unit is used to store a priority list of numbers for each contact and automatically try to dial multiple phone numbers in sequence when dialing until the call is connected or the list is exhausted; The message delivery unit is used to automatically select or use a pre-recorded message from a preset message template library and send it to the contact via at least one of SMS, instant messaging application or system-level push when all preset phone numbers are unreachable. The communication auditing unit is used to digitally sign the communication information using the electronic signature private key stored in the terminal after each call or message is sent, generate a unique communication signature Seal_ID and store it. The security management unit is used to interface with the five-digit authentication system to verify user identity and manage encryption keys.
2. The special population assisted communication system according to claim 1, characterized in that: The local voiceprint comparison engine adopts a voiceprint feature comparison method based on Gaussian mixture model-general background model; Specifically, for the i-th voice sample of each contact, the Mel-frequency cepstral coefficient feature vector sequence X is extracted. i ={x1, x2, ..., x T The archived model λ is derived adaptively from the general background model through maximum a posteriori probability. i The real-time voice feature vector sequence Y = {y1, y2, ..., y} input by the user. L Match score with the contact archive model In the formula, p is the probability density function, p(y) t |λ i ) indicates that in model λ i The observed eigenvector y t The probability density value; When the matching score exceeds a preset first threshold θ p When a match is found, it is determined that the contact is successfully matched; when multiple contacts have matching scores exceeding θ p When selecting a match, the highest matching score is chosen if the difference between it and the second-highest matching score is greater than a preset second threshold θ. m The contacts are used as the matching results.
3. The special population assisted communication system according to claim 1, characterized in that, The one-click multi-connection unit uses a combined judgment method based on call status code and duration to determine if a call is not connected. The system monitors call status changes through the TelephonyManager interface of the terminal operating system. When the status changes from DIALING to IDLE, the system obtains the call termination reason code. If the reason code is any one of CALL_FAIL_UNOBTAINABLE_NUMBER, CALL_FAIL_BUSY, CALL_FAIL_NO_ANSWER, CALL_FAIL_UNREACHABLE, or CALL_FAIL_TIMEOUT, or if the ringing continues for more than the preset ringing timeout threshold T, ring If the call never goes to OFFHOOK mode, it is determined that the current number is not connected. After determining that the call was not connected, the failure reason code of the current number is recorded, and the system automatically switches to the next number in the priority list to continue trying until a number is connected or the entire list has been tried. Before switching to the next number, a voice-synthesized announcement is made to inform the user of the current contact status.
4. The special population assisted communication system according to claim 1, characterized in that, The guaranteed message delivery unit includes a message delivery state machine, whose state transition logic is as follows: The initial state is IDLE. When the one-click multi-connection unit reports that all number attempts have failed, the state changes to MSG_PREPARE. The system selects the message template with the highest matching degree from the preset message template library according to the context. The message template includes replaceable variable placeholders, including the current time, the list of numbers tried, and the user's preset nickname. When the status changes to MSG_SENDING, the system attempts to send messages according to the following preset message channel priority list in sequence: the first channel is SMS, sent through SmsManager; if SMS sending fails, it switches to the second channel, instant messaging application messages, sent through the application's public interface; if the instant messaging application is not installed or sending fails, it switches to the third channel, system-level notification push. Once a message is successfully sent through any channel, the status changes to MSG_SENT, a message signature is generated, and the channel type and timestamp of the successful message transmission are recorded. If all channels fail to send the message, the status will change to MSG_FAILED, the system will generate a failure report stamp, and notify the user and / or their family members through the system notification. The message template library supports remote updates by family members. Updates take effect after being authenticated by the Five Data Package, and new templates can only be written to the local template library after being signed with an electronic seal.
5. The special population assisted communication system according to claim 1, characterized in that, The communication signature Seal_ID generated by the communication auditing unit adopts a timestamp-based notation method using blockchain light nodes, specifically including: After each call or message is sent, the system combines the signature data into a summary message M={user_id, contact_id, timestamp, trigger_type, confirm_type, call_duration / msg_hash, nonce}, where user_id is the unique identifier of the current user in the system, contact_id is the unique identifier of the contact of the other party, timestamp is the timestamp of the communication, trigger_type is the triggering method of this communication, confirm_type is the method of secondary confirmation, the call_duration / msg_hash field is filled with the hash value of the call duration or message content according to the communication type, and nonce is a randomly generated one-time random number; The digest information M is signed using the SM2 elliptic curve public key cryptography algorithm using the electronic signature private key to generate a signature value Sig. The hash value H(M) of the digest information M is packaged with the signature value Sig and submitted to the consortium blockchain network through a blockchain light node client to obtain the block height and transaction hash returned by the blockchain. The block height and transaction hash are associated with and stored with local signature data to form an immutable chain of communication evidence for subsequent tracing and auditing.
6. The special population assisted communication system according to claim 1, characterized in that, The specific method for the security management unit to interface with the five-digit authentication system is as follows: When a user activates the system for the first time, an identity binding request is initiated to the five-digit authentication server by reading the unique device identifier in the terminal security chip and the user's five-digit QR code. After the five-digit authentication server verifies the user's identity, it returns the identity credentials, including the user's public key certificate, to the terminal security chip for storage. The security chip meets the national cryptographic level 2 or above security standards. All contact photos are stored after being symmetrically encrypted with AES-256-GCM. The encryption key is protected by a key derived from inside the security chip, and no external process can directly read the plaintext photo data. Before each dialing attempt, the system verifies whether the current user's identity matches the bound identity stored in the security chip. Only after successful verification can the contact's photo and voice sample library be loaded.
7. The special population auxiliary communication system according to any one of claims 1 to 6, characterized in that, The custom physical triggering unit also includes an action feature template learning module, which performs the following steps: During the configuration phase, three-axis time-series data from the terminal's built-in accelerometer and / or gyroscope are continuously collected when the user performs the same target-triggered action multiple times, constructing an original action dataset A={a1, a2, ..., a N }, where, a n =(t n , acc X (t n ), acc Y (t n ), acc Z (t n ), gyr X (t n ), gyr Y (t n ), gyr Z (t n )) In the formula, acc X (t n ) indicates that at t n The acceleration value in the X-axis direction at time acc Y (t n ) indicates that at t n The acceleration value in the Y-axis direction at any given time, acc Z (t n ) indicates that at t n The acceleration value in the Z-axis direction at any given time; gyr X (t n ) indicates that at t n The angular velocity value of rotation about the X-axis at any given time, gyr Y (t n ) indicates that at t n The angular velocity value of rotation about the Y-axis at any given time, gyr Z (t n ) indicates that at t n The angular velocity value of rotation around the Z-axis at any given time; The three-axis time series data is segmented using a sliding window. Within each window of length W, the statistical feature vector F = [mean, std, max, min, range, zero-crossing rate, energy] is calculated, where mean is the mean of each axis data, std is the standard deviation of each axis data, max is the maximum value of each axis data, min is the minimum value of each axis data, range is the range of each axis data, zero-crossing rate is the rate of zero crossing of each axis data, and energy is the energy value of each axis data. The dynamic time warping algorithm is used to calculate the current input action sequence Q={q1, q2, ..., q M } and the stored corresponding contact C k Template action sequence P k ={p1, p2, ..., p K Optimal path distance between} In the formula, π is the monotonic alignment path and d is the Euclidean distance; When the dynamic time warp distance is less than the preset third threshold θ d At that time, it is determined that the current input action is related to contact C. k The trigger action was successfully matched, triggering a call to contact C. k .
8. The special population auxiliary communication system according to any one of claims 1 to 6, characterized in that, The secondary confirmation unit employs a time-window-based multimodal confirmation signal fusion and determination method, specifically including: After triggering the dialing, a process of length T is started. w The confirmation time window simultaneously monitors the following three confirmation signal sources: touch screen tap signal, accelerometer shaking signal, and specific voice confirmation word signal collected by microphone. Calculate the confidence score for each channel, and the confidence score c of the knock signal. tap The confidence level of the shaking signal, c, is calculated by matching the detected number of taps with the preset number of confirmed taps in the range [0,1]. shake The range ∈[0,1] is obtained by mapping the dynamic time warping distance between the shaking pattern and the preset shaking template to the [0,1] interval, and the confidence level c of the voice confirmation word is obtained. voice ∈[0,1] is obtained through the posterior probability output by the local speech keyword recognition model; Calculate the fusion confirmation score C=w1·c tap +w2·c shake +w3·c voice In the formula, w1+w2+w3=1 and each weight is dynamically adjusted according to the user's preset enable status. The weight of the disabled signal source is reset to zero, and the remaining weights are normalized proportionally. When the fusion confirmation score C exceeds the preset fourth threshold θ c If the second confirmation is successful, the dialing will proceed; otherwise, it will proceed within the time window T. w The dialing will be automatically canceled after the timeout period.
9. A special population-assisted communication method based on multimodal perception and a local trusted execution environment, applied to the special population-assisted communication system according to any one of claims 1 to 8, comprising the following steps: Step S1: In response to the first mode switching command, switch the terminal address book interface to photo grid mode, where each contact entry is associated with a photo, and clicking on the photo will directly trigger the dialing of the corresponding phone number; Step S2: Establish a voice sample voiceprint feature vector library for each contact in the terminal's local security chip. Use the Gaussian mixture model-general background model method to compare the user's real-time voice input with the local sample library. If a match is successful, a call is triggered. Step S3: Provide a custom trigger event setting interface, receive the mapping relationship between the non-voice physical trigger action set by the user and the contact, wherein the physical trigger action includes at least one of the following: screen tap count, device shake mode, touch area coordinates, external switch signal, and match and identify the current input action and the template action through a dynamic time warping algorithm; Step S4: After any triggering event is activated, the identity information of the contact to be dialed is broadcast through voice synthesis. Multimodal confirmation signals are collected and fused within a preset time window, and the fused confirmation score is calculated. The actual dialing is only performed after the fused confirmation score exceeds a preset threshold. Step S5: In response to dialing trigger, try to dial multiple phone numbers in sequence according to the preset priority list, and determine whether each number is connected in real time by listening to the call status code and ringing duration, until the connection is completed or the list is exhausted; Step S6: If all preset phone numbers cannot be reached, the message delivery process will be automatically activated. The preset message will be sent to the contact via SMS, instant messaging application or system-level push according to the message channel priority, and the sending status of each channel will be recorded. Step S7: After each call or message is sent, use the electronic signature private key stored in the terminal security chip to digitally sign the communication information, generate a unique communication signature Seal_ID, and timestamp it through a blockchain light node and store it in the local secure storage area.
10. The assisted communication method for special populations according to claim 9, characterized in that, The screen tap count recognition in step S3 employs a time-series pulse detection algorithm based on an adaptive threshold, specifically including: The touchscreen sensor monitors finger touch and release events, recording a timestamp t for each tap event. k and its pressure value p k This constitutes a knocking event flow. E={e1, e2, …, e n } In the formula, e k =(t k , p k ); Using a sliding time window T sliding Group the event stream and count the number of taps cnt within the window. When the time interval between two consecutive taps is Δt=t... k+1 -t k Less than the first interval threshold τ short Jitters determined to be from the same tap are merged; when Δt is greater than the second interval threshold τ, the jitters are merged. long The time is considered the end of a valid tap; A pressure adaptive threshold mechanism is introduced, which collects the pressure value set P of n taps by the user during the configuration phase. calib ={p1, p2,..., p n }, calculate its mean μ p and standard deviation σ p The effective striking condition is the pressure value p. k >μ p -α·σ p Where α is a preset sensitivity coefficient; At the end of the detection window, the number of valid taps counted is matched with the number of taps preset for each contact in the configuration table, and the contact that matches perfectly and has the smallest deviation in pressure characteristics is selected as the target contact.