Universal game platform system based on multi-level language unit and implementation method thereof

By combining multi-level language unit annotation tasks with multiple game platforms, users contribute voice data to acquire virtual resources during entertainment, solving the problems of low voice assistant recognition rate and high cost of online games, and achieving low-cost and efficient data collection and improved user stickiness.

CN121846683APending Publication Date: 2026-04-14田忍峰
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-16
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, voice assistants have a low recognition rate for personal accents, dialects, and special pronunciation habits. Furthermore, online game platforms have high user acquisition costs and retention rates depend on game content updates, lacking a virtuous cycle that effectively combines voice data collection with user entertainment.

Method used

By using multi-level language unit annotation tasks and combining multiple game platforms, users contribute voice data during entertainment, acquire virtual resources, and consume these resources in different games, forming a closed-loop economic system.

Benefits of technology

It achieves full-coverage voice data collection, acquires high-value data at low cost, improves user stickiness, provides multiple monetization channels, forms a self-circulating economy, optimizes AI models, and attracts more users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121846683A_ABST
    Figure CN121846683A_ABST
Patent Text Reader

Abstract

The invention discloses a universal game platform system based on a multi-level language unit and an implementation method of the universal game platform system, and relates to the technical field of voice recognition technologies and game platforms. The system comprises a voice data acquisition module used for acquiring user voice and decomposing the user voice into language units, and the language units comprise at least one of phoneme-level units, vocabulary-level units or sentence-level units. And the virtual resource allocation module is used for allocating virtual resource units according to the user labeling quantity. And the multi-game access module supports at least more than two different types of games. And the resource consumption module allows the user to consume virtual resource units in the game. And the resource backflow module is used for monitoring the resource quantity and guiding user backflow. The game currency is obtained through voice annotation, the game currency is used for closed-loop design of multi-game consumption, the technical effects of obtaining high-value voice data at low cost, high user stickiness and multi-realization are achieved, and the method can be widely applied to the fields of voice AI data collection and game platform operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of speech recognition technology and game platform technology, and in particular to a system and method for acquiring virtual resources through multi-level language unit annotation tasks and circulating them among multiple games. Background Technology

[0002] With the rapid development of artificial intelligence technology, voice assistants have become a standard feature of smart devices. However, the recognition rate of general speech models for individual accents, dialects, and special pronunciation habits is still relatively low, mainly due to the lack of targeted personalized training data.

[0003] On the other hand, online gaming platforms typically use a model where users purchase in-game currency, requiring them to pay to gain access to the game. This model suffers from high user acquisition costs and a high retention rate that depends on game content updates.

[0004] Existing technologies include paid acquisition of voice annotation data through crowdsourcing platforms, but these methods are costly; game platforms also use ad viewing to award game sessions, but this results in a poor user experience due to users passively receiving ads. How to integrate voice data collection with user entertainment needs to create a virtuous cycle is a pressing issue that needs to be addressed in the current technological field. Summary of the Invention

[0005] Purpose of the invention

[0006] This invention aims to provide a general game platform system based on multi-level language units and its implementation method. By combining multi-level speech annotation tasks with multiple game platforms, users can contribute speech data while having fun, achieving a win-win situation for users, platforms, and AI models.

[0007] Technical solution

[0008] The technical solution provided by this invention is as follows:

[0009] A general-purpose game platform system based on multi-level language units, comprising:

[0010] The voice data acquisition module is used to acquire the user's original voice input and decompose the voice input into language units, wherein the language units include at least one of phoneme-level units, vocabulary-level units, or sentence-level units.

[0011] A virtual resource allocation module, connected to the voice data acquisition module, is used to allocate a corresponding number of virtual resource units to the user account based on the number of annotations the user makes on the language units.

[0012] A multi-game access module is used to support at least two different types of game applications to access the platform.

[0013] The resource consumption module is connected to the multi-game access module and the user account, and is used to allow users to consume the virtual resource units in the different types of game applications.

[0014] The resource return module is used to monitor the number of virtual resource units in a user's account. When the number falls below a preset threshold, the user is guided back to the voice data acquisition module to continue acquiring virtual resource units.

[0015] Beneficial effects

[0016] Compared with the prior art, the present invention has the following beneficial effects:

[0017] First, comprehensive protection. It covers three levels of language units: phoneme, vocabulary, and sentence, ensuring that all speech input falls within the protection scope.

[0018] Second, high-value data can be acquired at low cost. Users voluntarily contribute voice annotation data to gain game access, and the platform does not need to pay any cash costs.

[0019] Third, high user engagement. A variety of game options cater to different user preferences, and virtual resources become universal assets across games.

[0020] Fourth, a self-circulating economic system. This forms a virtuous cycle where labor generates resources, resources are used for entertainment, and resources are depleted for further labor.

[0021] Fifth, multiple monetization channels. Advertising revenue, in-app purchase revenue, and data asset revenue are combined.

[0022] Sixth, the data flywheel effect. As the number of users and the amount of data grow, AI models are continuously optimized, attracting more users to join. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the module structure of the system in an embodiment of the present invention. In the figure, 101 is the voice data acquisition module, 102 is the virtual resource allocation module, 103 is the user account, 104 is the resource consumption module, 105 is the multi-game access module, and 106 is the resource return module.

[0024] Figure 2 This is a schematic diagram of the speech data acquisition and offloading process in an embodiment of the present invention. In the diagram, 201 is the step of acquiring user speech, 202 is the step of multi-level language unit decomposition, 203 is the step of AI pre-annotation and confidence calculation, 204 is the offloading decision judgment node, 205 is the automatic input channel, 206 is the self-calibration channel, 207 is the crowdsourced calibration channel, 208 is the mapping table update step, and 209 is the step of obtaining virtual resources.

[0025] Figure 3This is a schematic diagram illustrating the cycle of virtual resource flow and game consumption in an embodiment of the present invention. In the diagram, 301 represents voice annotation, 302 represents obtaining virtual resources, 303 represents game consumption, 304 represents a virtual resource shortage state, and 305 represents ad viewing relief.

[0026] Figure 4 This is a schematic diagram of the user behavior path in an embodiment of the present invention. In the diagram, 401 is the APP download node, 402 is the new user guide node, 403 is the voice recognition task completion node, 404 is the voice coin acquisition node, 405 is the game consumption node, and 406 is the voice recognition return node. Arrows indicate the user behavior path, and dashed arrows indicate the return loop path. Detailed Implementation

[0027] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0028] Example 1: System Architecture

[0029] like Figure 1 As shown, the present invention provides a general game platform system based on multi-level language units, comprising:

[0030] The 101 speech data acquisition module is used to collect user speech and decompose it into language units. This module has a built-in multi-level language unit recognition model, which can segment continuous speech streams into phoneme-level units, word-level units, or sentence-level units.

[0031] The 102 virtual resource allocation module, connected to the 101 voice data acquisition module, allocates a corresponding number of virtual resource units (hereinafter referred to as voice coins) to the 103 user account based on the number of language units annotated by the user. For example, each phoneme unit annotated adds 1 voice coin to the account; each vocabulary unit annotated adds 3 voice coins; and each sentence unit annotated adds 10 voice coins.

[0032] The platform includes 105 game integration modules, supporting various games such as Dou Dizhu, billiards, and match-3 games. Game developers can connect to the platform via API to share the user pool and voice currency economy system.

[0033] The 104 resource consumption module connects to the 105 multi-game access module and the 103 user account. Users consume voice coins when entering a game. For example, in Dou Dizhu (a popular Chinese card game), the normal game consumes 10 coins per game, while the expert game consumes 50 coins per game; in billiards, the game consumes 20 coins per game; and in Candy Crush, each level consumes 1 coin.

[0034] The 106 resource feedback module connects to the 103 user account to monitor the user's voice coin balance. When the balance falls below a preset threshold, a push notification guides the user back to the 101 voice data collection module to continue annotation; it also provides a relief option of earning 20 coins by watching ads.

[0035] Example 2: Multi-level Speech Data Processing Flow

[0036] As Figure 2 shown, the speech data acquisition and shunt processing includes the following steps:

[0037] 201 Step of collecting user speech, for example, the user reads aloud "Today is Monday".

[0038] 202 Step of multi-level language unit decomposition, the language unit decomposition unit segments the speech into multi-level units. Phoneme-level units include jin, tian, shi, xing, qi, yi; word-level units include today, is, Monday; sentence-level unit includes "Today is Monday".

[0039] 203 Step of AI pre-annotation and confidence calculation, the AI pre-annotation unit generates candidate words and confidence levels for each language unit. The phoneme-level unit jin corresponds to "今" with a confidence level of 95% and "金" with 3%; the word-level unit "今天是" corresponds to "今天是" with a confidence level of 98%; the sentence-level unit "今天是星期一" corresponds to a complete match with a confidence level of 99%.

[0040] 204 Shunt decision-making judgment node, the shunt decision-making unit makes a judgment based on the confidence level and language unit level. Phonemes with a high confidence level greater than 80% enter the 205 automatic entry channel; words with a medium confidence level between 30% and 80% enter the 206 self-calibration channel; sentences with a low confidence level less than 30% enter the 207 crowdsourcing calibration channel.

[0041] 208 Step of updating the mapping table, the mapping table is updated after the user completes the annotation.

[0042] 209 Step of obtaining virtual resources, the virtual resource allocation module allocates voice coins to the 103 user account according to the annotation unit level and quantity.

[0043] Example 3: Virtual Resource Economic Cycle

[0044] As Figure 3 shown, the virtual resource flow forms a closed loop.

[0045] 301 Speech annotation, the user obtains voice coins through speech annotation, and the voice coins are deposited into 302 to obtain virtual resources, that is, the 103 user account. The user selects a game in the game hall and enters 303 to play the game and consume. After the voice coins are consumed, the system detects the 304 state of insufficient virtual resources and guides the user to return to 301 for speech annotation to continue obtaining, forming a cycle.

[0046] When a user runs out of voice coins, a relief mechanism is triggered. The relief mechanism involves watching an ad for 30 seconds to earn 20 voice coins. The platform receives advertising revenue, and the user is allowed to continue playing the game.

[0047] Example 4: User Behavior Path Example

[0048] like Figure 4 As shown, the behavior path of a new user is as follows.

[0049] 401. Download App Node: The user downloads the app. 402. New User Guide Node: The user enters the new user guide. 403. Complete Voice Recognition Task Node: The system prompts the user to complete ten voice recognition tasks to earn 100 voice coins; the user completes the voice recognition tasks. 404. Earn Voice Coins Node: The user earns 100 voice coins, which are deposited into the user's account (103). 405. Play Game Node: The user enters the game lobby, selects Dou Dizhu (a popular Chinese card game), and plays five rounds, consuming 50 voice coins. 406. Return to Voice Recognition Node: After playing five rounds, the user has 50 voice coins remaining. The system prompts the user that they don't have enough voice coins to complete more voice recognition tasks; the user returns to the 403 "Complete Voice Recognition Task" node to continue earning voice coins.

[0050] After long-term use, users accumulate voice coins, ranks, friends, and skins, becoming active users of the platform and forming a network like... Figure 4 The reflux loop is indicated by the dashed arrow.

[0051] Example 5: Systematic Summary of Speech Phonemes

[0052] To comprehensively cover various pronunciation characteristics of users, this invention systematically summarizes speech phonemes within the 101 speech data acquisition module, forming a complete phoneme library. The phoneme library includes the following categories.

[0053] Classified by language type, phonemes include Mandarin, dialect phonemes, and foreign phonemes. Mandarin contains twenty-one initials, thirty-nine finals, and four tones, such as b, p, m, and f. Dialect phonemes refer to the unique pronunciations of various local dialects, such as the entering tone in Cantonese and the voiced consonants in Wu dialect. Foreign phonemes refer to common foreign phonemes from languages ​​like English and Japanese, such as the 'th' sound in English.

[0054] Based on pronunciation difficulty, phonemes are categorized into standard phonemes, ambiguous phonemes, and special phonemes. Standard phonemes refer to those that are clearly pronounced and easily identified, such as standard Mandarin pronunciation. Ambiguous phonemes refer to those that are unclear and easily confused, such as the inability to distinguish between retroflex and alveolar consonants, or between front and back nasal consonants. Special phonemes refer to pronunciations unique to an individual, such as accents or unclear articulation.

[0055] Phonemes are categorized by frequency of use, including high-frequency, mid-frequency, and low-frequency phonemes. High-frequency phonemes are those that appear frequently in everyday language, such as n, i, and h. Mid-frequency phonemes are those that appear occasionally, such as zh, ch, and sh. Low-frequency phonemes are those that appear very rarely, such as the sounds of uncommon characters or ancient pronunciations.

[0056] The phoneme library is dynamically updated. The system continuously expands and optimizes the phoneme library based on user annotation data and crowdsourcing results collected during the 208 mapping table update process. Adding a phoneme means automatically adding it when a phoneme not present in the library is detected. Phoneme merging means merging multiple phonemes consistently annotated as the same character by users. Phoneme splitting means splitting and refining the same phoneme when it is annotated as different characters by users.

[0057] Example 6: Sentence-level calibration application scenario

[0058] Sentence-level unit calibration is particularly suitable for the following scenarios, where data is collected via the 101 speech data acquisition module and stored via the 208 mapping table update step.

[0059] Commonly used commands are calibrated, allowing users to calibrate high-frequency commands such as turning on the air conditioner and navigating home, ensuring accurate responses from the voice assistant.

[0060] Personalized quotes record users' frequently used catchphrases and common expressions.

[0061] The whole-sentence dictation function allows users to read the entire text aloud, and the system completes the calibration in one go.

[0062] Dialect sentence data collection is used for dialect preservation and research.

Claims

1. A universal game platform system based on multi-level language units, characterized in that, include: The voice data acquisition module is used to acquire the user's original voice input and decompose the voice input into language units, wherein the language units include at least one of phoneme-level units, vocabulary-level units, or sentence-level units. A virtual resource allocation module, connected to the voice data acquisition module, is used to allocate a corresponding number of virtual resource units to the user account based on the number of annotations the user makes on the language units. A multi-game access module is used to support at least two different types of game applications to access the platform. The resource consumption module is connected to the multi-game access module and the user account, and is used to allow users to consume the virtual resource units in the different types of game applications. The resource return module is used to monitor the number of virtual resource units in a user's account. When the number falls below a preset threshold, the user is guided back to the voice data acquisition module to continue acquiring virtual resource units.

2. The system according to claim 1, characterized in that, The phoneme-level unit includes phoneme units or phoneme combinations, and the phoneme unit is the smallest auditory discrimination unit; the vocabulary-level unit includes words, phrases or technical terms; the sentence-level unit includes complete sentences, commonly used instructions or user-defined sentences.

3. The system according to claim 1, characterized in that, The voice data acquisition module further includes: Language unit decomposition unit is used to divide speech signals into phoneme-level units, vocabulary-level units, or sentence-level units; AI pre-annotation unit is used to match candidate characters for the language unit and generate confidence scores; The triage decision unit is used to triage the language units to an automatic input channel, a self-calibration channel, or a crowdsourced calibration channel based on the confidence score and the language unit level.

4. The system according to claim 1, characterized in that, The virtual resource allocation module dynamically adjusts the allocation quantity based on the language unit level and annotation quality: A base number of virtual resources are allocated to phoneme-level units; Vocabulary-level units are allocated more virtual resources than the basic amount; Sentence-level units are allocated the highest number of virtual resources; Additional reward resources will be allocated to units with high confidence, units with bounties, or units that have been verified multiple times.

5. The system according to claim 1, characterized in that, The resource consumption module allows users to use virtual resource units to place bets, purchase game room entry qualifications, and redeem in-game items or skins during game matches.

6. The system according to claim 1, characterized in that, The resource return module includes an incentivized video ad unit that provides an option to watch ads to obtain temporary resource units when a user's virtual resource units fall below a preset threshold.

7. The system according to claim 1, characterized in that, The multi-game access module supports at least two of the following game types: card and board games, casual competitive games, and social gaming games.

8. A method for implementing a general game platform based on multi-level language units, characterized in that, Includes the following steps: Voice data acquisition steps: Acquire the user's raw voice input and decompose it into language units, wherein the language units include at least one of phoneme-level units, vocabulary-level units, or sentence-level units; Virtual resource allocation steps: Based on the number of annotations the user has made to the language units, allocate a corresponding number of virtual resource units to the user account; Game integration steps: Provide at least two different types of game applications for users to choose from; Resource consumption steps: Allow users to consume the virtual resource units in the different types of game applications; Resource return step: Monitor the number of virtual resource units in the user's account. When it falls below a preset threshold, guide the user back to the voice data acquisition step to continue acquiring virtual resource units.

9. The method according to claim 8, characterized in that, The voice data acquisition step further includes: Speech signals are segmented into phoneme-level units, vocabulary-level units, or sentence-level units; Match candidate characters for the language unit and generate a confidence score; Based on the confidence score and language unit level, the language units are distributed to automatic input, self-calibration, or crowdsourced calibration channels. Based on user-annotated data and crowdsourcing results, we continuously expand and optimize the phoneme library, including adding new phonemes, merging phonemes, or splitting phonemes.

10. A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method as claimed in claim 8 or 9.