Decision support methods, devices, equipment, media, and smart glasses
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-23
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]本发明提供一种辅助决策方法、装置、设备、介质和智能眼镜,用以解决现有技术中依赖用户自行决策准确率和效率较低的缺陷
[0038]本发明提供的辅助决策方法、装置、设备、介质和智能眼镜,基于当前视野图像进行场景意图解析,得到目标意图,从而可以获取目标意图对应的辅助信息以及生成携带有目标意图、辅助信息以及当前视野图像的提示语句,从而辅助决策模型能够基于提示语句准确且快速得到辅助决策结果,避免传统方法中依赖用户自行决策导致决策偏差且效率较低的问题。
Smart Images

Figure CN117453039B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus, device, medium, and smart glasses for assisting decision-making. Background Technology
[0002] AR glasses (Augmented Reality Glasses) are smart wearable devices that use augmented reality technology to display virtual information into the user's field of vision, thereby expanding the user's perception and interaction capabilities.
[0003] Existing AR glasses mainly display virtual information into the user's field of vision through augmented reality technology. After seeing the corresponding virtual information through AR glasses, users need to rely on their own knowledge to make judgments and decisions. However, in some cases, users have limited knowledge, which makes it impossible for them to make accurate decisions. Furthermore, relying on users to make decisions on their own is inefficient and results in a poor user experience. Summary of the Invention
[0004] This invention provides a method, apparatus, device, medium, and smart glasses for assisting decision-making, in order to address the shortcomings of existing technologies that rely on users to make decisions themselves, resulting in low accuracy and efficiency.
[0005] This invention provides a decision support method, comprising:
[0006] Capture the user's current field of view image;
[0007] Based on the current field-of-view image, scene intent is parsed to obtain the target intent;
[0008] Obtain the auxiliary information corresponding to the target intent, and generate a prompt statement carrying the target intent, the auxiliary information, and the current field of view image;
[0009] The prompt statement is sent to the auxiliary decision-making model so that the auxiliary decision-making model can make auxiliary decisions based on the prompt statement and obtain auxiliary decision results;
[0010] Receive and display the auxiliary decision-making results returned by the auxiliary decision-making model.
[0011] According to a decision-making assistance method provided by the present invention, the step of parsing scene intent based on the current field-of-view image to obtain the target intent includes at least one of the following:
[0012] Based on the semantic information of the current field-of-view image, scene intent is parsed to obtain the target intent;
[0013] The target intent is obtained by parsing the user's voice command corresponding to the current field of view image.
[0014] The target intent is obtained by parsing the scene intent based on the user's current state corresponding to the current field of view image.
[0015] According to a decision-making assistance method provided by the present invention, the step of determining the prompt statement includes:
[0016] Determine the prompt statement template corresponding to the target intent;
[0017] Based on the auxiliary information and the current field of view image, the prompt statement template is filled in to obtain the prompt statement.
[0018] According to a decision-making assistance method provided by the present invention, the step of determining the prompt statement includes:
[0019] Determine the prompt statement template corresponding to the target intent;
[0020] Based on the auxiliary information and the current field of view image, the prompt statement template is filled in to obtain the prompt statement.
[0021] According to a decision-making assistance method provided by the present invention, the acquisition of the user's current field of view image includes:
[0022] Upon detecting a collection command, the user's current field of view image is collected; the collection command includes at least one of voice commands, gesture commands, and control commands.
[0023] If the file size of the current field-of-view image exceeds a threshold, the current field-of-view image is compressed.
[0024] According to a decision support method provided by the present invention, the step of sending the prompt statement to a decision support model, so that the decision support model performs decision support based on the prompt statement and obtains a decision support result, includes:
[0025] The prompt statement is sent to the auxiliary decision-making model, so that the auxiliary decision-making model can translate the text information in the current field of view image and obtain the auxiliary decision result based on the text translation result and the prompt statement.
[0026] According to a decision support method provided by the present invention, the step of sending the prompt statement to a decision support model, so that the decision support model performs decision support based on the prompt statement and obtains a decision support result, includes:
[0027] The prompt statement is sent to the auxiliary decision-making model, so that the auxiliary decision-making model can translate the text information in the current field of view image and obtain the auxiliary decision result based on the text translation result and the prompt statement.
[0028] The present invention also provides a decision-making aid device, comprising:
[0029] The acquisition unit is used to acquire the user's current field of view image;
[0030] The parsing unit is used to perform scene intent parsing based on the current field-of-view image to obtain the target intent;
[0031] The acquisition unit is used to acquire auxiliary information corresponding to the target intent and generate a prompt statement carrying the target intent, the auxiliary information and the current field of view image;
[0032] The sending unit is used to send the prompt statement to the auxiliary decision-making model, so that the auxiliary decision-making model can make auxiliary decisions based on the prompt statement and obtain auxiliary decision results;
[0033] The receiving unit is used to receive and display the auxiliary decision-making results returned by the auxiliary decision-making model.
[0034] The present invention also provides smart glasses, including: the decision-making assistance device as described above.
[0035] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the above-described auxiliary decision-making methods.
[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the decision support method as described above.
[0037] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the above-described auxiliary decision-making methods.
[0038] The auxiliary decision-making method, device, equipment, medium, and smart glasses provided by this invention analyze scene intent based on the current field-of-view image to obtain the target intent. This allows the acquisition of auxiliary information corresponding to the target intent and the generation of prompt statements carrying the target intent, auxiliary information, and the current field-of-view image. As a result, the auxiliary decision-making model can accurately and quickly obtain auxiliary decision-making results based on the prompt statements, avoiding the problems of decision bias and low efficiency caused by relying on user self-decision in traditional methods. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0040] Figure 1 This is a flowchart illustrating the decision support method provided by the present invention;
[0041] Figure 2 This is a flowchart illustrating the method for determining prompt statements provided by the present invention;
[0042] Figure 3 This is a flowchart illustrating an implementation of step 110 in the decision support method provided by the present invention.
[0043] Figure 4 This is a schematic diagram of the auxiliary decision-making device provided by the present invention;
[0044] Figure 5 This is a schematic diagram of the auxiliary decision-making interaction based on smart glasses provided by the present invention;
[0045] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0047] Existing AR glasses mainly display virtual information into the user's field of vision through augmented reality technology. After seeing the corresponding virtual information through AR glasses, users need to rely on their own knowledge to make judgments and decisions. However, in some cases, users have limited knowledge, which makes it impossible for them to make accurate decisions. Furthermore, relying on users to make decisions on their own is inefficient and results in a poor user experience.
[0048] For example, at a traffic light intersection, a user needs to use the traffic light information displayed in AR glasses to determine whether it is permissible to turn based on their knowledge of traffic rules. If the user's knowledge of traffic rules is not comprehensive enough, they are prone to making incorrect decisions.
[0049] In response, this invention provides a decision-making support method. Figure 1This is a flowchart illustrating the decision support method provided by the present invention, as shown below. Figure 1 As shown, this method can be applied to smart glasses and includes the following steps:
[0050] Step 110: Acquire the user's current field of view image.
[0051] Specifically, the current field of view image can be understood as the visual information that can be observed at present. For example, when a user is wearing smart glasses, the current field of view image can be understood as the real-time image that the smart glasses can currently capture.
[0052] The current field of view image can be acquired in real time or triggered by user voice commands. For example, at a traffic light intersection, a user wearing smart glasses looks at the traffic light and issues a voice command, "Can I turn right?" The smart glasses acquire the current field of view image upon receiving the voice command. Additionally, the current field of view image can also be acquired based on user gesture commands. For instance, the smart glasses acquire the current field of view image when they detect a preset gesture (such as a specified gesture swiping across the front of the glasses, a tapping motion, etc.).
[0053] Step 120: Perform scene intent parsing based on the current field-of-view image to obtain the target intent.
[0054] Specifically, scene intent parsing can be achieved by extracting the corresponding scene intent from the current field of view image; the extracted scene intent is also the target intent. Optionally, the target object in the current field of view image can be identified, and the target intent can be determined by combining the target object's contextual information (such as environmental information, background information, etc.).
[0055] For example, if the current field of view image is a "traffic light image", and the target object in the current field of view image is identified as a traffic light, and the environmental information of the target object is a traffic light intersection, then based on the target object and the environmental information of the target object, the target intent can be determined to be "intersection navigation".
[0056] Step 130: Obtain the auxiliary information corresponding to the target intent, and generate a prompt statement carrying the target intent, auxiliary information, and current field of view image.
[0057] Specifically, the auxiliary information corresponding to the target intent can be understood as information related to the target intent. For example, if the target intent is "intersection navigation", the corresponding auxiliary information may include current location information, current environment information, navigation map information, current traffic information, etc.
[0058] Furthermore, the prompt statement can be the text obtained by filling the prompt statement template with the target intent, auxiliary information, and current visual field image. The prompt statement can also be understood as the text obtained by normalizing the target intent, auxiliary information, and current visual field image into a standard input format according to the prompt statement template. Different target intents typically correspond to different prompt statement templates. The prompt statement can be a prompt text.
[0059] For example, a sample prompt might be: "You are an AR assistant. Based on the following current field-of-view image information and auxiliary information, please determine the user's current scene and some key information within that scene, including textual auxiliary information. Combine this with the user's gestures or voice commands to determine the user's intent. Finally, you need to combine the auxiliary information and provide a suitable decision suggestion for the user. User gesture information: XXX, User voice command: XXX, Auxiliary information: XXX, Current field-of-view image information: XXX."
[0060] Step 140: Send the prompt statement to the auxiliary decision-making model so that the auxiliary decision-making model can make auxiliary decisions based on the prompt statement and obtain the auxiliary decision-making result.
[0061] In some specific implementations, the prompt statement can be input into a pre-built auxiliary decision model, which then makes an auxiliary decision based on the prompt statement to obtain the auxiliary decision result. The auxiliary decision model is trained based on sample prompt statements and corresponding sample auxiliary decision results. The auxiliary decision model can be built based on pre-trained models such as BERT (Bidirectional Encoder Representations from Transformers), XLNet (extreme Multi-label Learning Network), ROBERTa (Robustly Optimized BERT approach), and T5 (Text-to-Text Transfer Transformer).
[0062] Here, the decision support model can also be a large-scale model deployed in a chatbot with human-like characteristics, such as the IFlytek Spark model. Chatbots equipped with this decision support model can engage in dialogue with users by understanding and learning human language. Furthermore, they can interact with users based on the context of the dialogue and possess truly human-like communication abilities. In addition, they also possess human abilities such as editing, translation, and searching.
[0063] Step 150: Receive and display the auxiliary decision-making results returned by the auxiliary decision-making model.
[0064] Specifically, after receiving the auxiliary decision-making results returned by the auxiliary decision-making model, the auxiliary decision-making results can be displayed in voice form or in text form. This embodiment of the invention does not specifically limit the specific form of the auxiliary decision-making results.
[0065] The auxiliary decision-making method provided in this embodiment of the invention performs scene intent parsing based on the current field-of-view image to obtain the target intent. This allows the acquisition of auxiliary information corresponding to the target intent and the generation of prompt statements carrying the target intent, auxiliary information, and the current field-of-view image. As a result, the auxiliary decision-making model can accurately and quickly obtain auxiliary decision-making results based on the prompt statements, avoiding the problems of decision bias and low efficiency caused by relying on user self-decision in traditional methods.
[0066] Based on the above embodiments, scene intent parsing is performed on the current field-of-view image to obtain the target intent, including at least one of the following:
[0067] Scene intent is parsed based on semantic information of the current field of view image to obtain the target intent;
[0068] The target intent is obtained by parsing the user's voice command corresponding to the current field of view image.
[0069] The target intent is obtained by parsing the user's current state based on the current field of view image.
[0070] Specifically, the semantic information of the current field-of-view image can be understood as the semantic information of objects, scenes, events, etc., contained in the image. Based on the semantic information of the current field-of-view image, scene intent analysis is performed to obtain the target intent.
[0071] For example, if the current field of view image is a "traffic light image", then by combining the semantic information in the current field of view image, the target intent can be determined to be "intersection navigation".
[0072] Furthermore, the user voice command corresponding to the current field of view image can be the voice command issued by the user to the current field of view image when it is being acquired. The user voice command can contain the user's intent, i.e. the target intent. Then, based on the user voice command corresponding to the current field of view image, scene intent analysis can be performed to obtain the target intent.
[0073] For example, if a user's voice command is "Can I turn right at this intersection?", the corresponding target intent is "intersection navigation".
[0074] The current user state corresponding to the current field-of-view image can be the user's current state at the time the current field-of-view image is acquired. The user's current state can include the user's selected mode (such as sports, leisure, etc.), the user's current real-time status information (such as location, speed, heart rate, etc.), and the user's personalized information (such as age, hobbies, etc.). Based on the user's current state, the user's intention, i.e., the target intention, can be inferred.
[0075] For example, if we know from the user's current status that the user will navigate to the swimming pool every day at 19:00, then if we collect the current field of vision image at 19:00, we can determine that the target's intention is "navigation".
[0076] Based on any of the above embodiments Figure 2 This is a flowchart illustrating the method for determining prompt statements provided by the present invention, as shown below. Figure 2 As shown, the methods for determining the prompt statement include:
[0077] Step 210: Determine the prompt statement template corresponding to the target intent;
[0078] Step 220: Based on auxiliary information and the current field of view image, fill in the prompt statement template to obtain the prompt statement.
[0079] Specifically, the prompt statement is the text obtained by filling the prompt statement template with auxiliary information and the current field of view image. The prompt statement can also be understood as the text obtained by shaping the auxiliary information and the current field of view image into standard input form according to the prompt statement template. The prompt statement can be a prompt text.
[0080] Furthermore, different target intentions correspond to different prompt statement templates. By parsing the scene intent based on the current field of view image, the target intent can be determined. Then, the corresponding prompt statement template can be determined based on the target intent. Finally, the prompt statement template can be filled in based on auxiliary information and the current field of view image to obtain the prompt statement.
[0081] For example, a sample prompt might be: "You are an AR assistant. Based on the following current field-of-view image information and auxiliary information, please determine the user's current scene and some key information within that scene, including textual auxiliary information. Combine this with the user's gestures or voice commands to determine the user's intent. Finally, you need to combine the auxiliary information and provide a suitable decision suggestion for the user. User gesture information: XXX, User voice command: XXX, Auxiliary information: XXX, Current field-of-view image information: XXX."
[0082] Based on any of the above embodiments, the prompt statement template is filled in based on auxiliary information and the current field of view image to obtain a prompt statement, including:
[0083] Fill the corresponding positions in the prompt statement template with auxiliary information, current field of view image and related information to obtain the prompt statement; the related information includes the user's voice command corresponding to the current field of view image and / or the user's current state corresponding to the current field of view image.
[0084] Specifically, the relevant information includes user voice commands corresponding to the current visual field image and / or the user's current state corresponding to the current visual field image. The user voice commands corresponding to the current visual field image can be voice commands issued by the user regarding the current visual field image during its acquisition. The user's current state corresponding to the current visual field image can be the user's current state at the time the current visual field image is acquired. The user's current state can include the user's selected mode (e.g., sports, leisure), the user's current real-time status information (e.g., location, speed, heart rate), and the user's personalized information (e.g., age, hobbies).
[0085] Based on auxiliary information and the current field of view image, relevant information is added to the prompt statement template, making the resulting prompt statement contain richer information, so that the auxiliary decision-making model can obtain auxiliary decision results more accurately based on the prompt statement.
[0086] Based on any of the above embodiments Figure 3 This is a flowchart illustrating an implementation method for step 110 in the decision support method provided by the present invention, as shown below. Figure 3 As shown, step 110 includes:
[0087] Step 111: Upon detecting a collection command, collect the user's current field of view image; the collection command includes at least one of voice commands, gesture commands, and control commands;
[0088] Step 112: If the file size of the current field of view image is greater than the threshold, compress the current field of view image.
[0089] Specifically, voice commands can be instructions issued by the user in the form of speech. For example, if a user sees an object in front of them, they can issue a voice command "What's in front of me?" Upon detecting this voice command, the current field of vision image is captured.
[0090] Furthermore, when a voice command is detected, intent analysis can be performed based on the voice information of the command to determine whether the user needs to have their current visual field image captured. For example, if the detected voice command is an exclamation from the user, "Ah, so beautiful," it can be determined that the user does not need to have their current visual field image captured for decision-making, and therefore, capturing the current visual field image is unnecessary.
[0091] Gesture commands can be instructions issued by the user through hand gestures. For example, if the user needs to capture the current field of vision image, they can move their palm across the front of their field of vision. If the gesture command is detected, the current field of vision image will be captured.
[0092] Control commands can be instructions issued by the user pressing a control. For example, if a user needs to capture the current field of view image, the user can press the corresponding control to issue the corresponding control command to capture the current field of view image.
[0093] In addition, if the file size of the current field of view image is larger than the threshold after the current field of view image is acquired, it indicates that the current field of view image is too large. At this time, the current field of view image can be compressed to save storage space and speed up the transmission.
[0094] Based on any of the above embodiments, step 140 includes:
[0095] The prompt statement is sent to the auxiliary decision-making model so that the auxiliary decision-making model can translate the text information in the current field of view image and obtain the auxiliary decision result based on the text translation result and the prompt statement.
[0096] Specifically, if there is text translation content in the current field of view (e.g., if there is English in the current field of view and it needs to be translated into Chinese), the prompt statement will indicate the text translation content that needs to be translated. This assists the decision-making model in translating the text information in the current field of view after receiving the prompt statement, and obtains the auxiliary decision result based on the text translation result and the prompt statement.
[0097] For example, if the current field of view image is the English sign "West Gate" for the west gate of the museum, the prompt will suggest that the text information in the current field of view image needs to be translated. Then, the auxiliary decision model will translate the text "West Gate" in the current field of view image to obtain the text translation result "West Gate". If the prompt also includes the user's voice command "How to get to the east gate", the auxiliary decision model will generate an auxiliary decision result to prompt the user "East Gate navigation" based on the prompt and the text translation result.
[0098] Based on any of the above embodiments, the auxiliary decision-making model is also used to determine the user's initial decision result based on the current field of view image and / or auxiliary information, and to return corresponding prompt information if the user's initial decision result is inconsistent with the auxiliary decision result.
[0099] Specifically, the user's initial decision can be understood as the decision made by the user based on their own knowledge. However, since users have limited knowledge, their initial decision may be unreasonable, requiring a prompt to the user. Therefore, this embodiment of the invention returns a corresponding prompt when the user's initial decision differs from the auxiliary decision, indicating that the initial decision may be unreasonable.
[0100] For example, if the auxiliary information shows that the user's current location is in a left-turn lane for motor vehicles, the user's initial decision can be determined as "turn left at the current intersection." However, the auxiliary decision-making model, based on the current intersection information in the auxiliary information, learns that the road is congested when turning left, which takes a long time. At this point, the auxiliary decision-making model may arrive at the decision result of "the road ahead is congested, it is not recommended to turn left at the current intersection," meaning the auxiliary decision result is inconsistent with the user's initial result. In this case, a prompt message "It is not recommended to turn left at the current intersection, the road ahead is congested" can be returned to remind the user to reconsider their decision.
[0101] The auxiliary decision-making device provided by the present invention is described below. The auxiliary decision-making device described below can be referred to in correspondence with the auxiliary decision-making method described above.
[0102] Based on any of the above embodiments Figure 4 This is a schematic diagram of the auxiliary decision-making device provided by the present invention, as shown below. Figure 4 As shown, the device includes:
[0103] Acquisition unit 410 is used to acquire the user's current field of view image;
[0104] The parsing unit 420 is used to perform scene intent parsing based on the current field-of-view image to obtain the target intent;
[0105] The acquisition unit 430 is used to acquire auxiliary information corresponding to the target intent and generate a prompt statement carrying the target intent, auxiliary information and the current field of view image;
[0106] The sending unit 440 is used to send the prompt statement to the auxiliary decision-making model so that the auxiliary decision-making model can make auxiliary decisions based on the prompt statement and obtain auxiliary decision results.
[0107] The receiving unit 450 is used to receive and display the auxiliary decision-making results returned by the auxiliary decision-making model.
[0108] Based on any of the above embodiments, scene intent parsing is performed on the current field-of-view image to obtain the target intent, including at least one of the following:
[0109] Scene intent is parsed based on semantic information of the current field of view image to obtain the target intent;
[0110] The target intent is obtained by parsing the user's voice command corresponding to the current field of view image.
[0111] The target intent is obtained by parsing the user's current state based on the current field of view image.
[0112] Based on any of the above embodiments, the steps for determining the prompt statement include:
[0113] Determine the prompt statement template corresponding to the target intent;
[0114] Based on auxiliary information and the current field of view image, the prompt statement template is filled in to obtain the prompt statement.
[0115] Based on any of the above embodiments, the prompt statement template is filled in based on auxiliary information and the current field of view image to obtain a prompt statement, including:
[0116] Fill the corresponding positions in the prompt statement template with auxiliary information, current field of view image and related information to obtain the prompt statement; the related information includes the user's voice command corresponding to the current field of view image and / or the user's current state corresponding to the current field of view image.
[0117] Based on any of the above embodiments, the user's current field of view image is acquired, including:
[0118] Upon detecting a collection command, the user's current field of view image is collected; the collection command includes at least one of voice commands, gesture commands, and control commands.
[0119] If the file size of the current field-of-view image exceeds a threshold, compress the current field-of-view image.
[0120] Based on any of the above embodiments, a prompt statement is sent to the decision support model so that the decision support model makes an auxiliary decision based on the prompt statement and obtains the auxiliary decision result, including:
[0121] The prompt statement is sent to the auxiliary decision-making model so that the auxiliary decision-making model can translate the text information in the current field of view image and obtain the auxiliary decision result based on the text translation result and the prompt statement.
[0122] Based on any of the above embodiments, the auxiliary decision-making model is also used to determine the user's initial decision result based on the current field of view image and / or auxiliary information, and to return corresponding prompt information if the user's initial decision result is inconsistent with the auxiliary decision result.
[0123] Based on any of the above embodiments, the present invention also provides smart glasses, including: the decision-making assistance device as described in any of the above embodiments.
[0124] As an optional embodiment, Figure 5 This is a schematic diagram of the auxiliary decision-making interaction based on smart glasses provided by the present invention, such as... Figure 5 As shown, when a user issues a voice command "Can we go now?", the smart glasses capture the current field of view image (traffic light image). At the same time, the smart glasses generate a prompt statement based on the current field of view image and the voice command and send it to the auxiliary decision-making model. The auxiliary decision-making model determines the auxiliary decision result based on the prompt statement and returns it, so the smart glasses display the auxiliary decision result "Yes".
[0125] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 6 As shown, the electronic device may include a processor 610, a memory 620, a communication interface 630, and a communication bus 640, wherein the processor 610, memory 620, and communication interface 630 communicate with each other through the communication bus 640. The processor 610 can call logical instructions in the memory 620 to execute an auxiliary decision-making method, which includes: acquiring a user's current field-of-view image; performing scene intent parsing based on the current field-of-view image to obtain a target intent; acquiring auxiliary information corresponding to the target intent and generating a prompt statement carrying the target intent, the auxiliary information, and the current field-of-view image; sending the prompt statement to an auxiliary decision-making model so that the auxiliary decision-making model performs auxiliary decision-making based on the prompt statement to obtain an auxiliary decision result; and receiving and displaying the auxiliary decision result returned by the auxiliary decision-making model.
[0126] Furthermore, the logical instructions in the aforementioned memory 620 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0127] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the auxiliary decision-making method provided by the above methods, the method comprising: acquiring a user's current field of view image; performing scene intent parsing based on the current field of view image to obtain a target intent; acquiring auxiliary information corresponding to the target intent, and generating a prompt statement carrying the target intent, the auxiliary information, and the current field of view image; sending the prompt statement to an auxiliary decision-making model, so that the auxiliary decision-making model performs auxiliary decision-making based on the prompt statement to obtain an auxiliary decision-making result; receiving and displaying the auxiliary decision-making result returned by the auxiliary decision-making model.
[0128] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned auxiliary decision-making methods. The method includes: acquiring a user's current field-of-view image; performing scene intent parsing based on the current field-of-view image to obtain a target intent; acquiring auxiliary information corresponding to the target intent and generating a prompt statement carrying the target intent, the auxiliary information, and the current field-of-view image; sending the prompt statement to an auxiliary decision-making model, so that the auxiliary decision-making model performs auxiliary decision-making based on the prompt statement to obtain an auxiliary decision result; and receiving and displaying the auxiliary decision result returned by the auxiliary decision-making model.
[0129] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0130] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A decision support method, characterized in that, include: Capture the user's current field of view image; Based on the current field-of-view image, scene intent is parsed to obtain the target intent; Obtain auxiliary information corresponding to the target intent, and generate a prompt statement carrying the target intent, the auxiliary information, and the current field of view image; the prompt statement is generated according to a prompt statement template, carrying text that includes the target intent, the auxiliary information, and the current field of view image; The prompt statement is sent to the auxiliary decision-making model so that the auxiliary decision-making model can make auxiliary decisions based on the prompt statement and obtain auxiliary decision results; Receive and display the auxiliary decision-making results returned by the auxiliary decision-making model.
2. The decision support method according to claim 1, characterized in that, The process of parsing scene intent based on the current field-of-view image to obtain the target intent includes at least one of the following: Based on the semantic information of the current field-of-view image, scene intent is parsed to obtain the target intent; The target intent is obtained by parsing the user's voice command corresponding to the current field of view image. The target intent is obtained by parsing the scene intent based on the user's current state corresponding to the current field of view image.
3. The decision support method according to claim 1, characterized in that, The steps for determining the prompt statement include: Determine the prompt statement template corresponding to the target intent; Based on the auxiliary information and the current field of view image, the prompt statement template is filled in to obtain the prompt statement.
4. The decision support method according to claim 3, characterized in that, The step of filling the prompt statement template with the auxiliary information and the current field of view image to obtain the prompt statement includes: The auxiliary information, the current field of view image, and related information are filled into the corresponding positions in the prompt statement template to obtain the prompt statement; the related information includes the user's voice command corresponding to the current field of view image and / or the user's current state corresponding to the current field of view image.
5. The decision support method according to any one of claims 1 to 4, characterized in that, The acquisition of the user's current field of view image includes: Upon detecting a collection command, the user's current field of view image is collected; the collection command includes at least one of voice commands, gesture commands, and control commands. If the file size of the current field-of-view image exceeds a threshold, the current field-of-view image is compressed.
6. The decision support method according to any one of claims 1 to 4, characterized in that, The step of sending the prompt statement to the decision support model, so that the decision support model can make an auxiliary decision based on the prompt statement and obtain an auxiliary decision result, includes: The prompt statement is sent to the auxiliary decision-making model, so that the auxiliary decision-making model can translate the text information in the current field of view image and obtain the auxiliary decision result based on the text translation result and the prompt statement.
7. The decision support method according to any one of claims 1 to 4, characterized in that, The auxiliary decision-making model is also used to determine the user's initial decision result based on the current field of view image and / or the auxiliary information, and to return corresponding prompt information if the user's initial decision result is inconsistent with the auxiliary decision result.
8. A decision-making auxiliary device, characterized in that, include: The acquisition unit is used to acquire the user's current field of view image; The parsing unit is used to perform scene intent parsing based on the current field-of-view image to obtain the target intent; The acquisition unit is used to acquire auxiliary information corresponding to the target intent and generate a prompt statement carrying the target intent, the auxiliary information and the current field of view image; the prompt statement is generated according to a prompt statement template, which is text carrying the target intent, the auxiliary information and the current field of view image. The sending unit is used to send the prompt statement to the auxiliary decision-making model, so that the auxiliary decision-making model can make auxiliary decisions based on the prompt statement and obtain auxiliary decision results; The receiving unit is used to receive and display the auxiliary decision-making results returned by the auxiliary decision-making model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the decision support method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the decision support method as described in any one of claims 1 to 7.
11. A type of smart glasses, characterized in that, include: The decision support device as described in claim 8.
Citation Information
Patent Citations
Image generation method and device, electronic equipment and storage medium
CN116580408A
Intelligent software testing
US20220058114A1