Audio data transmission and processing system based on proximity aware network
The audio data transmission and processing system based on proximity sensing networks solves the problem of data gaps in the recording data transmission and processing of TWS earphones, enabling seamless flow and intelligent processing of recording data, improving user experience and reducing costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN FENGHEYUAN TECH
- Filing Date
- 2026-06-01
- Publication Date
- 2026-06-30
AI Technical Summary
The existing TWS earphone recording function has gaps in the transmission and processing of recording data, resulting in a poor user experience, especially in meeting scenarios where the transmission and processing of recording data is disjointed and inefficient.
An audio data transmission and processing system based on a proximity sensing network is adopted. Through the automatic connection between the headphone compartment and the user terminal, Wi-Fi Aware technology is used to achieve seamless discovery and communication. Combined with a remote server, audio data is automatically analyzed and text is generated to establish an intelligent knowledge base to improve processing efficiency.
It enables seamless transfer of audio data, from audio capture to text display, significantly improving processing efficiency and user experience, reducing product material costs, and enhancing security and transmission reliability.
Smart Images

Figure CN122317569A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of wireless communication, and more specifically, to an audio data transmission and processing system based on a proximity sensing network. Background Technology
[0002] TWS earphones have become quite common in daily life, especially in the workplace where users may need to use earphones. Currently, the functions of earphones are constantly expanding to meet work requirements.
[0003] Meetings are an essential part of work, and TWS earbuds can record audio for easy note-taking. However, current technologies face several challenges in processing the recorded audio data. For example, the post-processing chain for audio data often suffers from significant gaps. Users who need to transcribe or summarize the recordings are typically limited by cumbersome manual processes, resulting in disjointed and inefficient audio data transmission and processing, leading to a poor user experience. Therefore, ensuring both reliable and intelligent transmission remains a crucial technological challenge. Summary of the Invention
[0004] In view of the above problems, this application proposes an audio data transmission and processing system based on a proximity sensing network, which can improve the coherence and efficiency of audio data transmission and processing, thereby improving the user experience.
[0005] This application provides an audio data transmission and processing system based on a proximity sensing network. The system includes: an earphone case configured to respond to a recording trigger operation, collect and store audio data, and automatically establish a communication connection with a user terminal when the proximity sensing network detects the user terminal; the user terminal configured to receive the audio data sent from the earphone case based on the proximity sensing network; and a remote server configured to communicate with the user terminal, receive the audio data forwarded from the user terminal, analyze the audio data, generate text information, and then feed the text information back to the user terminal for display.
[0006] Therefore, by utilizing a proximity sensing network, the headphone compartment can automatically establish a connection after discovering the user terminal, and the user terminal can automatically forward the audio data to the remote server, thus achieving a seamless flow from recording and acquisition to text display, significantly improving processing efficiency and user experience. Attached Figure Description
[0007] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments and drawings obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0008] Figure 1 This illustration shows a schematic diagram of an audio data transmission and processing system based on a proximity sensing network, as provided in an embodiment of this application.
[0009] Figure 2 This illustration shows a schematic diagram of another audio data transmission and processing system based on a proximity sensing network provided in an embodiment of this application.
[0010] Figure 3 A schematic diagram of the structure of an earphone compartment provided in an embodiment of this application is shown. Detailed Implementation
[0011] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0013] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0014] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0015] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0016] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0017] Please see Figure 1 and Figure 2 , Figure 1 This illustration shows a schematic diagram of an audio data transmission and processing system based on a proximity sensing network, according to an embodiment of this application. Figure 2 This illustration shows a schematic diagram of another audio data transmission and processing system based on a proximity sensing network, provided in an embodiment of this application. Figure 1 and Figure 2 As shown, the audio data transmission and processing system 100 based on a proximity sensing network includes an earphone compartment 110, a user terminal 120, and a remote server 130.
[0018] The earphone compartment 110 is configured to respond to recording trigger operations, collect and store audio data, and automatically establish a communication connection with the user terminal 120 when it is detected based on the proximity sensing network. The user terminal 120 is configured to receive audio data sent from the earphone compartment 110 based on the proximity sensing network. The remote server 130 is configured to communicate with the user terminal 120, receive audio data forwarded from the user terminal 120, analyze the audio data, generate text information, and feed the text information back to the user terminal 120 for display.
[0019] The earphone case 110 can be used for TWS (True Wireless Stereo) earphones that support recording function.
[0020] The recording trigger can be initiated by the user via a physical button on the earphone compartment 110 or by a control command sent through the accompanying app, which instructs the earphone compartment 110 to begin audio recording. After the recording trigger is activated, the earphone compartment 110 begins operation, converting the analog audio signal into digital audio data and temporarily storing it.
[0021] Proximity sensing network can refer to Wi-Fi Aware (NAN) technology. This is a proximity service protocol based on the Wi-Fi standard, allowing devices to directly discover and establish connections with each other without the need for traditional Wi-Fi access points or internet connections. This application utilizes proximity sensing network to achieve seamless automatic discovery and communication establishment between the headphone compartment 110 and the user terminal 120.
[0022] In one specific implementation, when pairing a certain earphone charging case 110 with a user terminal 120 for the first time, the initial pairing with the earphone charging case 110 needs to be performed through the accompanying APP on the user terminal 120. After the initial pairing is completed, if the earphone charging case 110 enters the vicinity of the user terminal 120, there is no need to operate the APP again, click the connection button, or perform any pairing operation. The earphone charging case 110 will automatically recognize the paired user terminal 120 using Wi-Fi Aware technology and immediately start the automatic push of recording files, thereby improving the convenience and efficiency of use and truly realizing a seamless experience of automatically pushing recording files when nearby.
[0023] In one specific implementation, if the earphone case 110 completes initial pairing with multiple user terminals 120, and the earphone case 110 simultaneously enters the vicinity of multiple user terminals 120, then the earphone case 110 establishes a connection with the user terminal 120 that has recently successfully transmitted among the multiple user terminals 120; or, the user selects one user terminal 120 from the multiple user terminals 120 to establish a connection; or, the earphone case 110 establishes a connection with the user terminal 120 with the strongest signal among the multiple user terminals 120; or, based on the current application scenario, it is determined which specific user terminal 120 the earphone case 110 will establish a connection with.
[0024] For example, multiple user terminals 120 include computers and mobile phones. If the current application scenario is a meeting mode, the headphone compartment 110 establishes a connection with the computer; if the current application scenario is a mobile scenario, the headphone compartment 110 establishes a connection with the mobile phone.
[0025] User terminal 120 can be a smartphone or computer device with a companion recording app installed. User terminal 120 is usually pre-paired with headphone compartment 110 and has Wi-Fi Aware enabled. When it receives audio data sent by headphone compartment 110, the app on user terminal 120 automatically parses the audio data and forwards it to remote server 130 via the Internet, while simultaneously receiving and displaying the final processing result sent from remote server 130.
[0026] The remote server 130 can be an AI computing platform deployed in the cloud. The remote server 130 is connected to the user terminal 120 via the Internet. The remote server 130 is configured to receive and forward audio data, convert the audio into audio text using speech recognition technology, and further generate meeting minutes or content summaries based on the audio text using natural language processing technology, and finally send the text information back to the user terminal 120.
[0027] In one specific implementation, the remote server 130 also includes a semantic analysis module, which performs in-depth analysis on the audio text and content summary to extract structured metadata as semantic tags and knowledge fingerprints for the recorded content. The structured metadata may include topics, core concepts, important entities (e.g., names, locations, technical terms), timestamps, and content categories (e.g., XX project meeting, XX group seminar).
[0028] After receiving the text information and structured metadata returned by the remote server 130, the APP associated with the user terminal 120 no longer simply displays it, but uses the structured metadata to build an intelligent knowledge base locally.
[0029] Specifically, the intelligent knowledge base will employ a graph database or semantic index structure to associate different audio files with their key concepts and entities. For example, if the concept of "XX Project Conference" appears in multiple conferences, the knowledge base will establish links between these lectures and the "XX Project Conference" concept. Simultaneously, the app will build an efficient full-text search index for all audio content and combine it with semantic tags for multi-dimensional indexing.
[0030] To ensure that audio data can be accurately captured and efficiently and securely transferred to the remote server 130, this application features a specific design for the hardware architecture of the headphone charging case 110. In a particular implementation, please refer to... Figure 3 , Figure 3 This application provides a schematic diagram of the structure of an earphone case according to an embodiment of the present application. Figure 3 As shown, the earphone case 110 includes a microphone module 111, a storage module 112, and a near-field wireless communication module 113, specifically: The microphone module 111 is configured to collect audio data in response to a recording trigger operation; the storage module 112 is configured to store the collected audio data; the near-field wireless communication module 113 is configured to transmit the stored audio data to the user terminal 120 based on the Wi-Fi Aware protocol; the earphone compartment 110 is also configured to control the storage module 112 to release the storage space corresponding to the audio data in response to a transmission completion confirmation command from the user terminal 120.
[0031] The microphone module 111 can employ a dual-microphone array design, utilizing differential amplification or beamforming technology to suppress ambient background noise and highlight the user's voice when acquiring audio data. The microphone module 111 not only picks up sound waves but also works with an audio codec to convert analog sound signals into digital audio streams, such as PCM format.
[0032] When the user performs a recording trigger operation, the main control chip of the earphone compartment 110 generates an interrupt signal, which wakes up the microphone module 111 in low power mode and makes it start continuous sampling at a preset sampling rate, which can be 16kHz or 48kHz.
[0033] The storage module 112 can be located inside the earphone compartment 110, rather than in the earphone itself, to reduce the size and power consumption of the earphone itself. The storage module 112 may include non-volatile memory, such as serial flash memory.
[0034] The main control chip of the earphone compartment 110 receives the digital audio stream acquired by the earphone compartment 110 itself through an internal interface and writes the digital audio stream into the flash memory in the form of a file stream. The internal interface can be I2S or UART.
[0035] The near-field wireless communication module 113 can be a functional module that supports the Wi-Fi Aware protocol. When the earphone case 110 is close to the user terminal 120, the near-field wireless communication module 113 will broadcast a service discovery frame. Once the paired user terminal 120 is detected, the near-field wireless communication module 113 will automatically establish a Wi-Fi direct link and use the high bandwidth of the Wi-Fi direct link to quickly push the audio data in the storage module 112 to the user terminal 120.
[0036] After the user terminal 120 successfully receives and verifies the audio data, the application of the user terminal 120 will send a transmission completion confirmation instruction to the headphone compartment 110 through the established Wi-Fi direct link. The transmission completion confirmation instruction usually includes the file identifier or transaction number of the successfully transmitted data, which is used to clearly inform the headphone compartment 110 which data has safely arrived at the remote server 130 or the user terminal 120 and can be cleared.
[0037] When the headphone compartment 110 receives the transmission completion confirmation command, the main control chip will parse the file identifier in the command, locate the corresponding audio file in the storage module 112, and call the deletion interface to mark the flash memory sector occupied by the corresponding audio data as free or erasable. This means that the headphone compartment 110 does not need to be equipped with a large-capacity storage chip to store historical recordings for a long time. It only needs to be equipped with a small-capacity storage chip to meet the needs of users for long-term cyclic use, thereby significantly reducing the material cost of the product.
[0038] Therefore, through the proximity sensing network, the earphone compartment 110 automatically establishes a communication connection and uploads data when it detects the user terminal 120, without requiring the user to manually pair or click to transfer. This seamless connection, combined with the aforementioned automatic deletion technology, allows the entire recording and transcription process to proceed without user intervention, greatly enhancing the user experience in office or meeting scenarios.
[0039] Furthermore, in some embodiments, the user terminal 120 is configured to: read the packet header information in the audio data, parse out the packet sequence number, total data length, and checksum; pre-allocate a buffer in memory that is appropriate for the total data length; write the payload of the audio data into the corresponding position in the buffer according to the packet sequence number, and perform integrity verification on the written payload using the checksum; if the verification fails, generate a retransmission request containing the packet sequence number and send it to the headphone compartment 110 through the proximity sensing network to instruct the headphone compartment 110 to retransmit the audio data corresponding to the packet sequence number; if the verification passes, in response to the cumulative length of the received packets reaching the total data length, determine that the audio data reception is complete, reassemble the data in the buffer, and output it to the remote server 130.
[0040] The packet header does not contain the actual audio waveform data, but rather metadata that instructs the user terminal 120 on how to process subsequent data. The packet header typically has a fixed length and format and is located at the very beginning of the data packet; the user terminal 120 parses the packet header first when reading data.
[0041] The packet sequence number is a unique, incrementing identifier assigned to each packet. Since audio data may be fragmented into multiple small packets during transmission through a proximity-aware network, and out-of-order delivery may occur, the packet sequence number is used to identify the logical order of the packet within the entire audio stream. User terminal 120 uses the packet sequence number to detect packet loss or out-of-order packets, thereby ensuring the continuity of audio reconstruction.
[0042] The total data length field can be the total number of bytes of complete audio data in the current transmission task. After reading the total data length field, the user terminal 120 can know the size of the entire recording file, thereby accurately calculating the required memory space and avoiding data truncation due to insufficient space or resource occupation due to wasted space.
[0043] The checksum can be a characteristic value calculated based on the payload of the data packet using a specific algorithm. For example, the specific algorithm could be Cyclic Redundancy Check (CRC), MD5, or a hash algorithm. User terminal 120 uses the same algorithm to calculate the checksum on the received data and compares the result with the checksum in the packet header. If the two do not match, it indicates that a bit error or corruption occurred during wireless transmission.
[0044] After receiving the audio data sent by the headphone compartment 110, the user terminal 120 first reads the header information of each data packet. Based on the total data length, the user terminal 120 pre-allocates a buffer of matching size in memory. Then, based on the data packet sequence number, the user terminal 120 writes the audio content of each data packet to the corresponding position in the buffer.
[0045] For example, if each data packet is 1KB, the data packet with sequence number 5 will be written starting from the 5KB mark. In this way, even if the data packets arrive out of order, they will be in their proper places and will not be disordered.
[0046] During the writing process, user terminal 120 recalculates the checksum using the same algorithm and compares it with the checksum in the packet header. If they match, the data is complete and error-free; if they do not match, the data has been corrupted during transmission.
[0047] If the verification fails, the user terminal 120 will immediately generate a retransmission request, including the sequence number of the corrupted data packet, and send it back to the headphone compartment 110 via the proximity sensing network. Upon receiving the request, the headphone compartment 110 will retrieve the packet from its local storage and retransmit it.
[0048] In one specific implementation, if the verification passes, the data segment is marked as valid; if the verification fails, the data segment is determined to be corrupted, and the location is marked as needing repair to prevent erroneous audio data from contaminating the final recording file.
[0049] When the total amount of data received by user terminal 120 reaches the total data length and all verifications pass, the entire audio data reception is considered complete. At this time, user terminal 120 concatenates the data in the buffer in sequence to restore the complete audio data, and then uploads it to remote server 130 for speech transcription.
[0050] Therefore, by using a pre-allocated buffer and a sequence-based writing scheme, the problem of out-of-order data during network transmission can be effectively addressed, ensuring that audio data is accurately stored in memory. Simultaneously, by combining a verification and retransmission scheme, data integrity verification and precise repair are achieved, guaranteeing that the audio file ultimately uploaded to the remote server 130 is complete and usable.
[0051] Based on the aforementioned solution ensuring reliable audio data transmission, in order to address the issue of how the earphone compartment 110 can accurately select the optimal connection target from multiple terminals in a complex environment, in some embodiments, the earphone compartment 110 is further configured to: when multiple user terminals 120 are detected based on a proximity sensing network, determine the connection evaluation parameters corresponding to each user terminal 120; determine the comprehensive score value corresponding to each user terminal 120 based on the connection evaluation parameters and preset weight values; determine the user terminal 120 corresponding to the highest comprehensive score value as the final connected user terminal 120, and automatically establish a communication connection with the final connected user terminal 120.
[0052] The connection evaluation parameters can be various metrics used to characterize the connection compatibility between user terminal 120 and headphone compartment 110, as well as the processing capabilities of user terminal 120. In one specific implementation, the connection evaluation parameters include at least the received signal strength, link quality, historical transmission success rate, timeliness factor, and user terminal 120 priority for each user terminal 120.
[0053] The received signal strength can be the RF signal power value measured by the earphone compartment 110 when receiving broadcast or detection signals from the user terminal 120. The received signal strength directly reflects the physical distance and path loss between the earphone compartment 110 and the user terminal 120. The closer the received signal strength value is to 0, the stronger the signal, the more stable the physical connection, and the greater the basic bandwidth potential for data transmission.
[0054] Link quality can be considered as an indicator of the signal purity and stability of the current communication link. Link quality not only considers signal strength but also takes into account factors such as environmental noise and co-channel interference. In a specific implementation, link quality can be characterized by signal-to-noise ratio, bit error rate, or packet loss rate.
[0055] For example, even if the received signal strength is high, if severe environmental interference leads to an excessively high bit error rate, the link quality score will decrease, thereby preventing the headphone compartment 110 from connecting to a user terminal 120 that appears to have a full signal but is actually unable to transmit data.
[0056] The historical transmission success rate is a reliability indicator derived from statistical analysis of interaction records between the headphone compartment 110 and the user terminal 120 over a preset time period. The historical transmission success rate reflects the stability of the user terminal 120 when processing audio data transmission tasks.
[0057] For example, if a user terminal 120 experiences multiple connection interruptions, data packet loss, or transcription and upload failures in the past week, and its historical transmission success rate is low, the headphone case 110 will lower the score of the user terminal 120 during the evaluation, and will tend to select the user terminal 120 with more stable past performance.
[0058] The timeliness factor can be a correction coefficient used to adjust the weight of historical data in the overall evaluation. Since the wireless environment and device status are dynamic, transmission records more recent than the current time better reflect the current status of the device.
[0059] The role of the timeliness factor is to give more reference value to recent interaction records while reducing the influence of older records. For example, although a user terminal 120 may have an average historical success rate, it may have performed well in recent connections. Under the influence of the timeliness factor, its overall score will be improved, thus adapting to the dynamic changes in device status.
[0060] The priority of user terminal 120 can be a basic weight value set based on the device type, functional attributes, or user preset preferences of user terminal 120. The priority of user terminal 120 reflects the differences in the adaptability of different user terminals 120 in audio data processing services.
[0061] For example, in the default system configuration, smartphones are usually set to the highest priority because of their excellent call and network upload capabilities; tablets are next; while smartwatches or IoT devices have relatively lower priority due to their small screens or limited processing power.
[0062] For example, if the current system is in meeting mode and the current time is during working hours on a weekday, then the priority of user terminal 120 on the tablet is 1, and the priority of user terminal 120 on the smartwatch is 0.2.
[0063] User terminal 120 priority ensures that, when physical indicators such as signal strength are similar, the headphone case 110 will prioritize connecting to a device with more powerful functions and that is more in line with the user's main operating habits.
[0064] The overall score can be a numerical result calculated by the main control chip of the earphone case 110 after normalizing the connection evaluation parameters based on a preset weighted algorithm. Specifically, in some embodiments, the step "determine the overall score corresponding to each user terminal 120 according to the connection evaluation parameters and preset weight values" may include: determining the overall score corresponding to each user terminal 120 according to the connection evaluation parameters, preset weight values, and preset comprehensive scoring function.
[0065] In one specific implementation, the preset comprehensive scoring function can be expressed as: in, For user terminals The corresponding overall score, For user terminals The corresponding received signal strength, For user terminals The corresponding link quality, For user terminals The corresponding historical transmission success rate For user terminals The corresponding time factor, User terminal Corresponding user terminal priority, , , , and For preset weight values, This indicates normalization to the interval between 0 and 1.
[0066] In one specific implementation method It can be 0.2. It can be 0.1. It can be 0.3. It can be 0.2. It can be 0.2. It is understood that this application does not limit... to The specific value to be taken.
[0067] In other embodiments, the earphone compartment 110 is configured to: when multiple user terminals 120 are detected based on a proximity sensing network, receive active status identifiers sent from the multiple user terminals 120; classify the user terminals 120 in an active state into a candidate connection set based on the active status identifiers; and determine the user terminal 120 corresponding to the highest comprehensive score value as the target user terminal 120 based on the connection evaluation parameters corresponding to each user terminal 120 in the candidate connection set, and automatically establish a communication connection with the target user terminal 120.
[0068] The earphone case 110 activates a proximity sensing network for broadcasting or listening. When multiple paired user terminals 120 are detected in the vicinity, to avoid connecting to devices that the user is not currently interested in, the earphone case 110 further obtains the activity status identifiers of each device. Among them, the uninterested devices can be backup devices placed in a bag or in a locked state.
[0069] The active status identifier reflects the current interaction status of the user terminal 120. For example, the headphone case 110 can send a status query request to each user terminal 120 via a Bluetooth Low Energy auxiliary link, or directly parse the status bits carried by the user terminal 120 in the broadcast packet.
[0070] For example, if a mobile phone is in an unlocked state with the screen on, or a computer is in an unlocked working state, the active status it sends is marked as active; conversely, if the user's device is in a state of screen off, screen locked, or sleep, it is marked as inactive.
[0071] The earphone case 110 removes all devices marked as inactive from the connection list and only adds devices marked as active to the candidate connection set, ensuring that the earphone case 110 only focuses on devices that the user is currently operating or that are within sight.
[0072] For one or more user terminals 120 that enter the candidate connection set, the headphone compartment 110 further calculates their comprehensive score. This calculation process can refer to the comprehensive score calculation process described above and will not be repeated here. Finally, the headphone compartment 110 compares the comprehensive scores of all users in the candidate set, and locks the user terminal 120 with the highest comprehensive score as the target user terminal 120. After determining the target user terminal 120, the headphone compartment 110 establishes a connection with the target user terminal 120.
[0073] In other embodiments, the earphone compartment 110 is configured to: when multiple user terminals 120 are detected based on a proximity sensing network, determine the current scene mode of the system; filter out the target user terminal 120 corresponding to the scene mode from among the multiple user terminals 120, and automatically establish a communication connection with the target user terminal 120.
[0074] Considering that the environment of the user terminal 120 around the user is dynamically changing during the use of the earphone case 110, in some other embodiments, the earphone case 110 is also configured to: determine a first comprehensive score value of the currently connected user terminal 120; determine the comprehensive score value with the highest value as a second comprehensive score value; if the second comprehensive score value is greater than the first comprehensive score value, then establish a communication connection with the user terminal corresponding to the second comprehensive score value, and then disconnect the currently connected user terminal 120.
[0075] The second comprehensive score can be the highest comprehensive score among all user terminals 120 in the current proximity sensing network, excluding the currently connected user terminal 120.
[0076] If the second comprehensive score is greater than the first comprehensive score, it indicates that a better transmission node has appeared in the environment than the currently connected user terminal 120. The headphone compartment 110 automatically triggers the switching process, first disconnecting the communication connection with the current user terminal 120 (i.e., the user terminal 120 corresponding to the first comprehensive score), and then initiating a connection request to the user terminal 120 corresponding to the second comprehensive score to establish a new communication link.
[0077] Furthermore, in actual audio data processing, audio data is often not a single data stream, but rather contains multiple files to be transmitted. For example, a high-fidelity original recording file, a compressed transcribed audio text, and related metadata configuration files, etc.
[0078] Once a communication link is established, if network bandwidth is limited or during connection switching, transmitting all data indiscriminately in sequence may cause critical data to be blocked by large volumes of non-critical data, thus affecting the user's real-time experience. Therefore, based on the determination of the target transmission link, in some implementations, the headphone compartment 110 is also configured to: determine the transmission priority parameter corresponding to each of the multiple files to be transmitted; determine the priority score corresponding to each file to be transmitted based on the transmission priority parameter and a preset priority evaluation index; and sort and transmit the multiple files to be transmitted according to the priority score.
[0079] Transmission priority parameters can be used to characterize the business-level attributes and urgency of files to be transmitted. In one specific implementation, transmission priority parameters include at least the recording time, user intent factor, and retry compensation factor for each file to be transmitted.
[0080] In some implementations, the more recent the recording time of the file to be transmitted, the higher its priority.
[0081] The user intent factor is a numerical indicator that represents the urgency or level of attention a user shows towards a file to be transferred. Understandably, the more frequently a user interacts with a file and the more direct their actions, the higher their user intent factor, thus receiving a higher transfer priority.
[0082] For example, if a user marks a file to be transferred as urgent, then the user intent factor corresponding to that file is 1; otherwise, it is 0. As another example, the smaller a file to be transferred, the higher its priority.
[0083] The retry compensation factor is a numerical indicator that represents the number of historical transmission failures of a file due to network fluctuations or transmission errors. By introducing a retry compensation factor, it is possible to prevent an important file from being stuck in a vicious cycle of transmission failure-retry-failure due to occasional network jitter, which would cause the file to be continuously interrupted by newly generated files, ultimately resulting in data backlog or loss.
[0084] For example, the more times a file to be transferred is marked as having failed to transfer, the larger its corresponding retry compensation factor will be.
[0085] Priority score is a numerical result calculated based on transmission priority parameters and preset priority evaluation indicators. Priority score serves as a quantitative ranking basis for the transmission order of files to be transmitted.
[0086] Specifically, in some implementations, the step "determine the priority score corresponding to each file to be transmitted based on the transmission priority parameter and the preset priority evaluation index" may include: determining the priority score corresponding to each file to be transmitted based on the transmission priority parameter, the preset priority evaluation index and the multi-dimensional priority scoring function.
[0087] In one specific implementation, the multidimensional priority scoring function can be expressed as: in, File to be transferred The corresponding priority score, For the current time, File to be transferred The corresponding recording time, File to be transferred Corresponding user intent factors File to be transferred The corresponding file size, File to be transferred The corresponding retry compensation factor, , , and These are preset priority evaluation indicators.
[0088] In one specific implementation method It can be 0.4. It can be 0.3. It can be 0.2. It can be 0.1. It is understood that this application does not limit... , , and The specific value to be taken.
[0089] It is worth noting that if the earphone compartment 110 is removed or the signal is interrupted during transmission, the entire audio data needs to be retransmitted upon reconnection, which is inefficient and wastes the earphone compartment 110's power. Therefore, in some implementations, during the transmission of each audio data item, the user terminal 120 periodically sends an acknowledgment frame to indicate the offset of the successfully received bytes. The earphone compartment 110 maintains a transmission status table that records the transmitted offset of each audio data item.
[0090] When the transmission is interrupted due to signal loss, movement of the earphone compartment 110, or other reasons, and the earphone compartment 110 re-establishes a connection with the user terminal 120, the earphone compartment 110 first queries the user terminal 120 for the received offset of each audio data, and then continues the transmission from the point of interruption, instead of restarting the transmission of the entire audio data, in order to improve transmission efficiency and reduce the power consumption of the earphone compartment 110.
[0091] In practical use, transmitting large volumes of audio data takes a long time, causing smaller volumes of audio data to be blocked. Therefore, for large volumes of audio data, typically exceeding 10MB, the headphone compartment 110 divides them into fixed-size blocks, such as 256KB. The headphone compartment 110 can transmit multiple different blocks of audio data simultaneously, i.e., concurrent transmission, to fully utilize the bandwidth of Wi-Fi Aware. The user terminal 120 is responsible for reassembling the blocks into complete audio data.
[0092] Block transmission can employ a sliding window mechanism, which involves sending N blocks consecutively, waiting for batch confirmations, and automatically retransmitting unconfirmed blocks.
[0093] Therefore, the earphone compartment 110 achieves intelligent sorting of files to be transmitted by calculating priority scores, ensuring that high-value and time-sensitive audio data can occupy the communication link first. Considering that audio data may involve user privacy and security, in some embodiments, the earphone compartment 110 is also configured to: generate a first encryption key for each of the multiple files to be transmitted; encrypt each file to be transmitted according to the first encryption key corresponding to each file to be transmitted, generating encrypted data corresponding to each file to be transmitted; encrypt the first encryption key corresponding to each file to be transmitted according to a preset shared key, obtaining a second encryption key corresponding to each file to be transmitted; and transmit the data packet containing the encrypted data and the second encryption key corresponding to each file to be transmitted to the user terminal 120 based on a proximity sensing network.
[0094] The first encryption key is a temporary symmetric key, such as an AES-128 or AES-256 key. For each file to be transmitted in the transmission queue, the headphone compartment 110 randomly generates a unique key, ensuring that even if an attacker cracks the first encryption key of a file to be transmitted, they cannot use the first encryption key to decrypt another file, greatly improving the system's resistance to attacks and its security.
[0095] The earphone compartment 110 uses the generated first encryption key to perform high-speed symmetric encryption on the file to be transmitted, transforming the original file into unreadable gibberish, i.e., encrypted data. Because symmetric encryption algorithms have low computational complexity and high speed, this method is very suitable for real-time encryption of large data files like audio files, without causing significant transmission delays.
[0096] In order to securely send the first encryption key to the user terminal 120, the earphone compartment 110 uses a preset shared key to encrypt the first encryption key to obtain the second encryption key. The preset shared key may be negotiated and securely stored in advance by the earphone compartment 110 and the user terminal 120 during a previous pairing or binding phase.
[0097] The earphone compartment 110 packages the encrypted data together with the second encryption key and sends it to the user terminal 120. After receiving the data packet, the user terminal 120 uses a locally stored preset shared key to decrypt the second encryption key, recovering the first encryption key. Then, it uses the recovered first encryption key to decrypt the encrypted data, ultimately obtaining the original file to be transmitted. This improves the security of the file to be transmitted between the earphone compartment 110 and the user terminal 120.
[0098] In some implementations, the earphone compartment 110 is further configured to: detect the current distance between itself and the currently connected target user terminal 120 based on a proximity sensing network; when the current distance is greater than a preset safe distance threshold, cache the device address and negotiation parameters of the target user terminal 120 in a historical connection list and disconnect the communication connection with the target user terminal 120; generate random padding data and write the random padding data into the storage space of the first encryption key associated with the current session in a volatile memory; when the target user terminal 120 is detected again based on the proximity sensing network, extract the access device address from the broadcast signal from the target user terminal 120; compare the access device address with the historical connection list, and when a match is successful, read the negotiation parameters of the target user terminal 120 from the historical connection list; based on the read negotiation parameters, send a connection restoration request to the target user terminal 120 to skip the service discovery and security negotiation phase and re-establish the communication connection with the target user terminal 120.
[0099] The earphone case 110 uses a proximity sensing network to monitor the physical distance between itself and the target user terminal 120 in real time. Once the distance exceeds a preset safe distance threshold, the earphone case 110 determines that the target user terminal 120 has left the safe area. The preset safe distance threshold can be 3 meters, 5 meters, or 5.1 meters, and this application does not limit the specific value of the preset safe distance threshold.
[0100] Before disconnecting the communication connection with the target user terminal 120, the headphone compartment 110 stores the device address and negotiation parameters of the target user terminal 120 in the historical connection list. The device address can be a MAC address, and the negotiation parameters can be channel information and capability sets.
[0101] After disconnecting the communication connection with the target user terminal 120, if the first encryption key, originally stored in memory (volatile memory), is only processed by deleting the pointer, its actual binary data may still remain on the physical storage medium, posing a risk that malicious software can recover it through memory forensics. To completely eliminate this risk, the earphone compartment 110 generates meaningless random padding data, such as all 0s, all 1s, or random numbers, directly overwriting the memory address storing the first encryption key. This ensures that the key is not only logically deleted but also physically completely eliminated, preventing the risk of key leakage after disconnection.
[0102] When the target user terminal 120 returns to the coverage area of the headphone compartment 110, the headphone compartment 110 will capture the broadcast signal emitted by the target user terminal 120. The headphone compartment 110 will extract the device address in the broadcast and search for it in the historical connection list. Once a match is found, the headphone compartment 110 will directly read the previously cached negotiation parameters and directly initiate a connection restoration request, skipping the intermediate negotiation steps, thereby greatly shortening the reconnection time and reducing power consumption during the interaction process.
[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An audio data transmission and processing system based on a proximity sensing network, characterized in that, The system includes: The earphone case is configured to respond to recording trigger operations, collect and store audio data, and automatically establish a communication connection with the user terminal when the user terminal is detected based on the proximity sensing network. The earphone compartment includes a microphone module, a storage module, and a near-field wireless communication module; The microphone module is configured to collect the audio data in response to the recording trigger operation; The storage module is configured to store the collected audio data; The near-field wireless communication module is configured to transmit the stored audio data to the user terminal based on the Wi-Fi Aware protocol; The earphone case is further configured to: when multiple user terminals are detected based on the proximity sensing network, determine the connection evaluation parameters corresponding to each of the multiple user terminals; determine the comprehensive score value corresponding to each user terminal according to the connection evaluation parameters and the preset weight value; determine the user terminal corresponding to the highest comprehensive score value as the final connected user terminal, and automatically establish a communication connection with the final connected user terminal; The user terminal is configured to receive the audio data sent from the earphone compartment based on the proximity sensing network; The earphone case is also configured to respond to a transmission completion confirmation command from the user terminal and control the storage module to release the storage space corresponding to the audio data. A remote server is configured to communicate with the user terminal, receive the audio data forwarded from the user terminal, analyze the audio data, generate text information, and then send the text information back to the user terminal for display.
2. The audio data transmission and processing system based on a proximity sensing network according to claim 1, characterized in that, The earphone compartment is also configured to: Determine the first overall score of the currently connected user terminal; The highest overall score value is determined as the second overall score value; If the second comprehensive score is greater than the first comprehensive score, a communication connection is established with the user terminal corresponding to the second comprehensive score, and then the currently connected user terminal is disconnected.
3. The audio data transmission and processing system based on a proximity sensing network according to claim 1, characterized in that, The step of determining the comprehensive score for each user terminal based on the connection evaluation parameters and preset weight values includes: Based on the connection evaluation parameters, the preset weight values, and the preset comprehensive scoring function, a comprehensive score value corresponding to each user terminal is determined; wherein, the connection evaluation parameters include at least the received signal strength, link quality, historical transmission success rate, timeliness factor, and user terminal priority corresponding to each user terminal; The preset comprehensive scoring function is expressed as follows: in, For user terminals The corresponding comprehensive score value, For user terminals The corresponding received signal strength, For user terminals The corresponding link quality, For user terminals The corresponding historical transmission success rate, For user terminals The corresponding time factor, For user terminals The corresponding user terminal priority, , , , and The preset weight value is used.
4. The audio data transmission and processing system based on a proximity sensing network according to claim 1, characterized in that, The audio data includes multiple files to be transmitted; the earphone compartment is further configured to: Determine the transmission priority parameter for each of the plurality of files to be transmitted; Based on the transmission priority parameters and the preset priority evaluation index, determine the priority score corresponding to each file to be transmitted; The multiple files to be transmitted are sorted and transmitted according to the priority score.
5. The audio data transmission and processing system based on a proximity sensing network according to claim 4, characterized in that, The step of determining the priority score for each file to be transmitted based on the transmission priority parameter and the preset priority evaluation index includes: Based on the transmission priority parameters, the preset priority evaluation index, and the multi-dimensional priority scoring function, the priority score corresponding to each file to be transmitted is determined; wherein, the transmission priority parameters include at least the recording time, user intent factor, and retry compensation factor corresponding to each file to be transmitted; The multidimensional priority scoring function is expressed as follows: in, File to be transferred The corresponding priority score, For the current time, File to be transferred The corresponding recording time, File to be transferred The corresponding user intent factor, File to be transferred The corresponding file size, File to be transferred The corresponding retry compensation factor, , , and The preset priority evaluation index is referred to as the index.
6. The audio data transmission and processing system based on a proximity sensing network according to claim 1, characterized in that, The audio data includes multiple files to be transmitted; the earphone compartment is further configured to: Generate a first encryption key for each of the plurality of files to be transmitted; Each file to be transmitted is encrypted using the first encryption key corresponding to each file to be transmitted, thereby generating encrypted data corresponding to each file to be transmitted. Based on a preset shared key, the first encryption key corresponding to each file to be transmitted is encrypted to obtain the second encryption key corresponding to each file to be transmitted. Based on the proximity sensing network, a data packet containing encrypted data corresponding to each file to be transmitted and a second encryption key corresponding to each file to be transmitted is transmitted to the user terminal.
7. The audio data transmission and processing system based on a proximity sensing network according to claim 6, characterized in that, The earphone compartment is also configured to: Based on the proximity sensing network, the current distance between the target user terminal and the currently connected target user terminal is detected; When the current distance is greater than the preset safe distance threshold, the device address and negotiation parameters of the target user terminal are cached in the historical connection list, and the communication connection with the target user terminal is disconnected. Generate random padding data and write the random padding data into the storage space of the first encryption key associated with the current session in the volatile memory; When the target user terminal is detected again based on the proximity sensing network, the access device address is extracted from the broadcast signal from the target user terminal; The access device address is compared with the historical connection list. When a match is found, the negotiation parameters of the target user terminal are read from the historical connection list. Based on the read negotiation parameters, a connection restoration request is sent to the target user terminal to skip the service discovery and security negotiation phases and re-establish a communication connection with the target user terminal.
8. The audio data transmission and processing system based on a proximity sensing network according to claim 1, characterized in that, The user terminal is configured as follows: Read the header information from the audio data and parse out the data packet sequence number, total data length, and checksum. Pre-allocate a buffer in memory that is appropriate for the total length of the data; According to the data packet sequence number, the payload of the audio data is written into the corresponding position of the buffer, and the integrity of the written payload is verified using the checksum. If the verification fails, a retransmission request containing the data packet sequence number is generated and sent to the headphone compartment through the proximity sensing network to instruct the headphone compartment to retransmit the audio data corresponding to the data packet sequence number. If the verification passes, in response to the cumulative length of the received data packets reaching the total data length, it is determined that the audio data reception is complete, and the data in the buffer is reassembled and output to the remote server.