Game assisting method and device
By recognizing user voice and combining it with game images, and using a strategy knowledge base and multimodal model to generate voice guidance, the inefficiency and inaccuracy of existing game assistance solutions are solved, providing efficient and accurate guidance without affecting the gaming experience.
Patent Information
- Application Number
- CN202511045365.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing game assistance methods are unable to provide efficient and effective guidance without affecting the user's gaming experience. In particular, traditional text/graphic guides, video guides, and general voice assistants are unable to understand game scenarios in real time and provide accurate suggestions.
By recognizing user voice and converting it into text, matching game images, and leveraging game strategy knowledge bases and large multimodal models to generate guiding text and provide real-time suggestions in the form of voice, it combines visual information for intelligent responses.
It achieves efficient guidance without affecting the immersion during the game, provides accurate game suggestions, and improves the efficiency and accuracy of game assistance.
Smart Images

Figure CN120532141B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a game assisting method and device. Background Art
[0002] Currently, when gamers encounter difficulties, they usually rely on the following methods to get help:
[0003] 1. Text / Graphic Guides: Players need to pause the game and search for text / graphic guides through a browser or mobile phone. This method seriously interrupts the game immersion. In complex scenes, static graphics cannot clearly convey the timing and spatial relationships of operations.
[0004] 2. Video walkthrough: Players watch pre-recorded videos. This method is more intuitive, but its drawback is that its linear playback makes it difficult for players to quickly locate the time point that perfectly matches their current predicament. Players need to repeatedly drag the progress bar, which is inefficient and lacks interactivity.
[0005] 3. General voice assistants: Existing voice assistants (such as Siri and Alexa) lack a deep understanding of specific game content and cannot recognize visual scenes within the game, so they cannot provide effective help that is strongly related to the current game status.
[0006] In view of this, how to provide a game assistance solution that does not affect the user's gaming experience while the user is playing the game and can efficiently and effectively guide the user has become a technical problem that needs to be solved urgently. Summary of the Invention
[0007] In response to the technical problems existing in the prior art, the embodiments of the present application provide a game assistance method and device.
[0008] In a first aspect, an embodiment of the present application provides a game assistance method, comprising:
[0009] When the user is playing the game, the user's voice is recognized and converted into text, and a game image corresponding to the text is matched and recorded as the target game image;
[0010] Using the target game image to search the preset game strategy knowledge base, the target game strategy information is obtained;
[0011] By submitting the target game image, text and target game strategy information to a large multimodal model for analysis, guiding text is generated, and the guiding text is synthesized into speech and played.
[0012] In a second aspect, an embodiment of the present application further provides a game assisting device, comprising:
[0013] A matching unit is used to recognize the user's voice during the game, convert the voice into text, and match the game image corresponding to the text, which is recorded as the target game image;
[0014] A retrieval unit, configured to retrieve a preset game strategy knowledge base using a target game image to obtain target game strategy information;
[0015] The synthesis unit is used to generate guidance text by submitting the target game image, text and target game strategy information to a large multimodal model for analysis, synthesizing the guidance text into speech and playing it.
[0016] The game assistance method and device provided in the embodiments of the present application can understand the player's game scene and voice problems in real time, dynamically call relevant strategy knowledge, and provide intelligent responses in the form of voice. In the entire solution, the user does not need to leave the game interface and can get help through natural voice dialogue, which does not destroy the game immersion. Moreover, because it can "see" the player's current game screen, combined with visual information and problems, the accuracy of the suggestions provided is far superior to traditional methods. That is, it provides a game assistance solution that does not affect the user's gaming experience while the user is playing the game and can efficiently and effectively guide the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A flowchart of an embodiment of a game assisting method provided in an embodiment of the present application;
[0018] Figure 2 This is a structural diagram of an embodiment of a game auxiliary device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.
[0020] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0021] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0022] Reference Figure 1 FIG. 1 is a flow chart of a game assisting method provided in an embodiment of the present application, wherein the game assisting method includes:
[0023] S10, while the user is playing the game, the user's voice is recognized, the voice is converted into text, and a game image corresponding to the text is matched and recorded as a target game image;
[0024] In this embodiment, it should be noted that the game can be a PC game, a mobile game, a cloud game, etc. While the user is playing the game, it is necessary to monitor the user's voice input in real time and record the screen image (i.e., the game image displayed on the user's terminal screen). When the user's voice input is monitored, the voice is converted into text, and the game image corresponding to the text is determined from the recorded screen image, and this game image is recorded as the target game image.
[0025] S11, using the target game image to search a preset game strategy knowledge base to obtain target game strategy information;
[0026] In this embodiment, it should be noted that after determining the target game image, a preset game strategy knowledge base can be searched to obtain game strategy information related to the target game image. This game strategy information is the target game strategy information. The game strategy knowledge base contains multiple game images and their associated game strategy information.
[0027] S12. Generate guidance text by submitting target game images, text, and target game strategy information to a large multimodal model for analysis, synthesize the guidance text into speech, and play it.
[0028] For example, let's assume a cloud game is being played on a mobile phone. This solution captures the user's voice and screenshots in real time. At some point, the user is unsure what to do next. They can say, "What should I do next?" This solution captures the user's voice, converts it into text, and identifies the screenshot (assuming it's image P) taken at the start of the voice as the target game image. Image P is then used to search the game strategy knowledge base, obtaining strategy information related to image P: W1: Walk to the edge of a cliff. There's a hidden platform below. Jump down to pick up item M, and W2: Note: Jumping from this angle requires a running start; otherwise, you'll fall to your death. These W1 and W2 represent the target game strategy information. Image P, the text "What should I do next," and W1 and W2 are then fed to the GPT-4o model for analysis. This generates and plays back instructional text, such as: You're standing on the edge of a cliff. According to the strategy, you can try a running start and then jump to the platform directly below, where there's a hidden item.
[0029] The game assistance method provided in the embodiment of the present application can understand the player's game scene and voice problems in real time, dynamically call relevant strategy knowledge, and provide intelligent responses in the form of voice. In the entire solution, the user does not need to leave the game interface and can get help through natural voice dialogue, which does not destroy the game immersion. Moreover, because it can "see" the player's current game screen, combined with visual information and problems, the accuracy of the suggestions provided is far superior to traditional methods. That is, it provides a game assistance solution that does not affect the user's gaming experience while the user is playing the game and is efficient and effective.
[0030] Based on the aforementioned method embodiment, before using the target game image to search the preset game strategy knowledge base, the method may further include:
[0031] Download game walkthrough videos from video platforms in batches and separate them into video frame sequences and audio tracks;
[0032] The audio track is converted into text with time information, and the text with time information is temporally associated with the video frames in the video frame sequence. A game strategy knowledge base is constructed based on the association results between the text and the video frames.
[0033] In this embodiment, it should be noted that the temporal association of the text with time information with the video frames in the video frame sequence is essentially intended to extract game images and their corresponding game strategy information, i.e., game strategy knowledge, from the game strategy video. After association, a game strategy knowledge base can be constructed based on the association results. Furthermore, to speed up retrieval, the game images and game strategy information in the game strategy knowledge base can be represented as vectors.
[0034] Based on the aforementioned method embodiment, temporally associating the text with time information with the video frames in the video frame sequence may include:
[0035] The text with time information is associated with at least one video frame within a first target time period, wherein the first target time period is a time period corresponding to the text with time information.
[0036] In this embodiment, it should be noted that the first target time period may be the time period to which the text with time information belongs, a time period before the time period to which the text with time information belongs, or a time period after the time period to which the text with time information belongs.
[0037] Based on the aforementioned method embodiment, the method of using the target game image to search a preset game strategy knowledge base to obtain the target game strategy information may include:
[0038] By searching the game strategy knowledge base, game strategy information corresponding to at least one game image in the game strategy knowledge base having a similarity with the target game image greater than a preset threshold is obtained as the target game strategy information.
[0039] In this embodiment, it should be noted that when searching the game strategy knowledge base, the similarity between the target game image and the game images in the game strategy knowledge base can be compared. The N (N is a positive integer) game images in the game strategy knowledge base that are most similar to the target game image can be found. At least one game image can be identified from these N game images, and the game strategy information corresponding to this at least one game image can be used as the target game strategy information. However, if the game images and game strategy information in the game strategy knowledge base are represented as vectors, the target game image must first be converted into a vector during the search, and then the converted vector can be used for the search.
[0040] Based on the aforementioned method embodiment, matching the game image corresponding to the text may include:
[0041] The screen image of the user playing the game is acquired in real time, and a video frame within a second target time period is used as a target game image, wherein the second target time period is a time period corresponding to the voice.
[0042] In this embodiment, it should be noted that the second target time period may be the time period to which the speech belongs, a time period before the time period to which the speech belongs, or a time period after the time period to which the speech belongs.
[0043] Reference Figure 2 FIG. 1 is a schematic diagram of the structure of a game assisting device provided in an embodiment of the present application, the device comprising:
[0044] The matching unit 20 is used to recognize the user's voice during the game, convert the voice into text, and match the game image corresponding to the text, which is recorded as the target game image;
[0045] A retrieval unit 21 is used to retrieve a preset game strategy knowledge base using the target game image to obtain target game strategy information;
[0046] The synthesis unit 22 is used to generate guiding text by submitting the target game image, text and target game strategy information to a large multimodal model for analysis, synthesizing the guiding text into speech and playing it.
[0047] The game assistance device provided in the embodiment of the present application can understand the player's game scene and voice problems in real time, dynamically call relevant strategy knowledge, and provide intelligent responses in the form of voice. In the entire solution, the user does not need to leave the game interface and can get help through natural voice dialogue, which does not destroy the game immersion. Moreover, because it can "see" the player's current game screen, combined with visual information and questions, the accuracy of the suggestions provided is far superior to traditional methods. That is, it provides a game assistance solution that does not affect the user's gaming experience while the user is playing the game and is efficient and effective.
[0048] The game assistance device provided in the embodiment of the present application has an implementation process that is consistent with the game assistance method provided in the embodiment of the present application, and the effects that can be achieved are also the same as those of the game assistance method provided in the embodiment of the present application, which will not be repeated here.
[0049] Based on the above device embodiment, the device may further include:
[0050] The construction unit is used to batch download game strategy videos of games from the video platform before the retrieval unit works, separate the game strategy videos into video frame sequences and audio tracks, convert the audio tracks into text with time information, and temporally associate the text with time information with the video frames in the video frame sequence, and build a game strategy knowledge base based on the association results between the text and the video frames.
[0051] Based on the aforementioned device embodiment, the construction unit can be used to:
[0052] The text with time information is associated with at least one video frame within a first target time period, wherein the first target time period is a time period corresponding to the text with time information.
[0053] Based on the aforementioned device embodiment, the retrieval unit may be used to:
[0054] By searching the game strategy knowledge base, game strategy information corresponding to at least one game image in the game strategy knowledge base having a similarity with the target game image greater than a preset threshold is obtained as the target game strategy information.
[0055] Based on the aforementioned device embodiment, the matching unit may be used to:
[0056] The screen image of the user playing the game is acquired in real time, and a video frame within a second target time period is used as a target game image, wherein the second target time period is a time period corresponding to the voice.
[0057] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A game assisting method, characterized in that: include: When the user is playing the game, the user's voice is recognized and converted into text, and a game image corresponding to the text is matched and recorded as the target game image; Using the target game image to search the preset game strategy knowledge base, the target game strategy information is obtained; By submitting the target game image, text and target game strategy information to a large multimodal model for analysis, guiding text is generated, and the guiding text is synthesized into speech and played.
2. The method according to claim 1, wherein Before using the target game image to retrieve a preset game strategy knowledge base, the method further includes: Download game walkthrough videos from video platforms in batches and separate them into video frame sequences and audio tracks; The audio track is converted into text with time information, and the text with time information is temporally associated with the video frames in the video frame sequence. A game strategy knowledge base is constructed based on the association results between the text and the video frames.
3. The method according to claim 2, wherein Temporally associating the text with time information with the video frames in the video frame sequence includes: The text with time information is associated with at least one video frame within a first target time period, wherein the first target time period is a time period corresponding to the text with time information.
4. The method according to any one of claims 1 to 3, wherein The method of using the target game image to search a preset game strategy knowledge base to obtain target game strategy information includes: By searching the game strategy knowledge base, game strategy information corresponding to at least one game image in the game strategy knowledge base having a similarity with the target game image greater than a preset threshold is obtained as the target game strategy information.
5. The method according to claim 1, wherein The matching of the game image corresponding to the text includes: The screen image of the user playing the game is acquired in real time, and a video frame within a second target time period is used as a target game image, wherein the second target time period is a time period corresponding to the voice.
6. A game assisting device, characterized in that: include: A matching unit is used to recognize the user's voice during the game, convert the voice into text, and match the game image corresponding to the text, which is recorded as the target game image; A retrieval unit, configured to retrieve a preset game strategy knowledge base using a target game image to obtain target game strategy information; The synthesis unit is used to generate guidance text by submitting the target game image, text and target game strategy information to a large multimodal model for analysis, synthesizing the guidance text into speech and playing it.
7. The device according to claim 6, characterized in that Also includes: The construction unit is used to batch download game strategy videos of games from the video platform before the retrieval unit works, separate the game strategy videos into video frame sequences and audio tracks, convert the audio tracks into text with time information, and temporally associate the text with time information with the video frames in the video frame sequence, and build a game strategy knowledge base based on the association results between the text and the video frames.
8. The device according to claim 7, wherein The building block is used to: The text with time information is associated with at least one video frame within a first target time period, wherein the first target time period is a time period corresponding to the text with time information.
9. The device according to any one of claims 6 to 8, characterized in that The retrieval unit is used to: By searching the game strategy knowledge base, game strategy information corresponding to at least one game image in the game strategy knowledge base having a similarity with the target game image greater than a preset threshold is obtained as the target game strategy information.
10. The device according to claim 6, wherein The matching unit is used to: The screen image of the user playing the game is acquired in real time, and a video frame within a second target time period is used as a target game image, wherein the second target time period is a time period corresponding to the voice.
Citation Information
Patent Citations
Data processing method, apparatus, device and storage medium
CN109063662A
Identifying player engagement to generate contextual game play assistance
CN112203732A