Ship CAD user-defined command accurate identification method

By establishing a hot word list in ship design and assigning weights to keywords, and combining multimodal input signals, the accuracy of speech recognition in ship modeling is solved, efficient and accurate modeling instruction recognition is achieved, and modeling efficiency and accuracy are improved.

CN120354474APending Publication Date: 2025-07-22中国船舶集团海舟系统技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510414155.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In ship structural design, traditional speech recognition technology is difficult to accurately identify ship professional terms and complex dimensional instructions, resulting in high modeling error rates and affecting modeling efficiency and accuracy.

Method used

The hot word recognition system is adopted to improve the accuracy of speech recognition by establishing a hot word list and assigning weights to keywords, and combining multimodal input signals, especially the recognition of ship professional vocabulary, size and materials.

Benefits of technology

It significantly improves the accuracy of speech recognition in ship modeling scenarios, reduces homophone misrecognition, ensures the precise execution of modeling instructions, and improves modeling efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354474A_ABST
    Figure CN120354474A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of CAD modeling, and particularly relates to a ship CAD user-defined command accurate identification method and system, and the method comprises the steps: building a hot word list based on ship design field terminologies, CAD operation instruction keywords, material specifications and size parameters; distributing a weight for each hot word in the hot word list; a user inputs a voice instruction through the acquisition equipment; a user voice instruction is received and decoded, when a hot word list recorded by a user contains content, a decoder tends to output professional vocabularies of the ship industry through weight distribution of hot words, and the instruction input by the user voice is recognized more rigorously and accurately. According to the method, hot word weights with different priorities are set for professional words such as professional vocabularies, common command words, parameter names, units or material specifications and the like in each design stage of a ship, so that the recognition error rate caused by homonyms in a ship modeling scene is remarkably increased, and a more accurate input text is provided for a CAD system to subsequently execute an intelligent modeling instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of CAD modeling, and specifically relates to a method and system for accurately identifying user-defined commands in ship CAD. Background Art

[0002] In the work of ship structure design, designers need to frequently call function options in CAD software and input a large number of parameters to complete modeling tasks such as hull, piping, and electrical in the overall design and detailed design stages. The traditional method of relying on the mouse and keyboard to complete modeling is cumbersome, requires high precision, is time-consuming, and prone to errors. To improve efficiency, using speech recognition technology combined with large language model parsing to input modeling instructions and parameters has become a new exploration direction for modeling interaction.

[0003] However, in the ship scenario, there are many professional terms and proper nouns, such as hatch, longitudinal girder, frame number, stiffener, etc., which are common and have specific meanings; in addition, there are a large number of dimensions, section specifications, material grades, and instruction formats, such as complex statements like "create a floor support with a width of 1200 millimeters at the foredeck". If only using a general language model for speech-to-text conversion, it is very easy to cause recognition errors due to homophony or low word frequency. For example, "hatch" may be misrecognized as "reference", and "frame number" may be misrecognized as "inner position", resulting in failed instruction parsing or incorrect modeling results in subsequent links.

[0004] To solve the above problems, the present invention proposes a method, system, and device for applying a keyword boosting system to ship voice command recognition. In the speech recognition stage, "hot words" are weighted or homophone replacement is prioritized for specific ship professional words, command keywords, and common dimension or material codes, which can effectively improve the recognition accuracy and ensure the reliability of the large language model's analysis of instruction intent and parameter reasoning in subsequent intelligent modeling links. Therefore, it involves a method for accurately identifying user-defined commands in ship CAD based on keyword recognition. Summary of the Invention

[0005] To solve the problems raised in the above background art, the present invention provides a method and system for accurately identifying user-defined commands in ship CAD, mainly used to enhance the accuracy of command recognition methods such as speech recognition in ship design, especially the recognition of keywords such as ship proper nouns, terms, dimension specifications, etc., and the recognition of instructions such as modeling and adjusting the perspective, so as to provide higher-quality input text for the intelligent modeling scenario of ship CAD structural parts.

[0006] To achieve the above object, the first aspect of the present invention provides a method for accurately identifying user-defined commands in ship CAD, including:

[0007] Step S100: Establish a hot word list based on professional terms in the ship design field, CAD operation instruction keywords, material specifications, and dimensional parameters;

[0008] Step S200: Assign weights to each hot word in the hot word list;

[0009] Step S300: The user enters a voice command through a collection device;

[0010] Step S400: Receive the user's voice command and decode it. When the content included in the hot word list is entered by the user, the decoder tends to output professional vocabulary in the ship industry through the weight assignment of the hot word.

[0011] As a technical solution of the present invention, the hot word list is stored in the JSON array format, and the fields include hot word text and hot word weight. When the voice recognition service is started, the configured JSON hot word list is read, and all hot words and their weights are parsed for the beam search algorithm of voice recognition decoding to match and read and adjust the weights.

[0012] As a technical solution of the present invention, a hot word enhancement mechanism is introduced in the beam search process of voice recognition decoding to maintain several candidate sequences of words with the highest probability. Each word has a corresponding weight score, and by adjusting the weight value of the hot word, the score of the word in the candidate sequence is changed;

[0013] By applying the hot word weight to the output result, the probability enhancement of specific words during inference is achieved; by giving an additional score bonus to the candidate words that appear in the hot word list, the probability of the candidate words in the decoding output is increased.

[0014] As a technical solution of the present invention, the weight value is mapped to a fixed probability bonus.

[0015] As a technical solution of the present invention, for the voice command entered by the user, an end-to-end ASR model is used for voice text conversion; it also includes:

[0016] Capture the user's gesture actions, use the OpenPose algorithm to identify the gesture type and align it with the voice command timing;

[0017] Obtain the user's line of sight focus area and lock the operation object to the model component within the current window;

[0018] Construct a multimodal instruction parsing engine, and integrate voice, gesture, and eye movement signals into a structured instruction through a timestamp synchronization mechanism.

[0019] As a technical solution of the present invention, the operation state of the CAD software is obtained in real time, including the current editing layer, the activated tool type, and the model component attributes, and the hot word weight is dynamically adjusted according to the operation state.

[0020] As a technical solution of the present invention, in step S100, the method for updating the hot word list includes:

[0021] When the user corrects the recognition error, the system records the error cases and generates a training data set, and uses an algorithm to optimize the hot word weights. The weight adjustment formula is as follows:

[0022]

[0023] Where α is the learning rate, R is the actual recognition accuracy rate, is the expected value, Sim is the semantic similarity function, and β is the correction coefficient;

[0024] When the user inputs an uncollected word, the system calculates its semantic similarity with the hot word library through the word vector model, and uses it as a candidate word for the user to confirm and then add it to the hot word list.

[0025] As a technical solution of the present invention, it also includes analyzing the user's historical operation sequence, predicting the user's future n-step instruction input, and preloading relevant hot words into the buffer area.

[0026] As a technical solution of the present invention, for obtaining the user's line of sight focus area and locking the operation object to the model component within the current window, it includes the combination and binding of the eye movement focus area and the voice command object. Specifically:

[0027] If the user mentions the ship structure A in the voice command and the eye movement focus is in the area of structure A, the operation object is automatically locked to component A;

[0028] If the eye movement focus area does not match the position of the structure A mentioned in the voice command, a confirmation mechanism is triggered, and the components corresponding to the candidate words are provided for the user to select.

[0029] The second aspect of the present invention proposes a precise recognition system for user-defined commands in ship CAD, including:

[0030] A hot word list construction module for establishing a hot word list based on professional terms in the ship design field, CAD operation instruction keywords, material specifications, and dimension parameters;

[0031] A weight management module for assigning weights to each hot word in the hot word list;

[0032] An acquisition module for the user to input voice commands through an acquisition device;

[0033] An output module for receiving the user's voice commands and decoding them. When the content included in the hot word list is input by the user, the decoder tends to output professional terms in the ship industry through the weight assignment of the hot words.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] The present invention performs more rigorous and accurate recognition of the instructions input by the user's voice. By setting "hot word weights" with different priorities for professional terms, common command words, parameter names, units, or material specifications, etc. in each design stage of the ship, the recognition error rate caused by homophones in the ship modeling scenario is significantly reduced, providing more accurate input text for the subsequent execution of intelligent modeling instructions by the CAD system. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic flow diagram of the method of the present invention;

[0037] Figure 2 It is a judgment diagram of the combination and binding of the eye movement focus area and the voice command object in the present invention;

[0038] Figure 3 It is a schematic diagram of the fusion of multi-modal input signals in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0040] As Figures 1-3 shown. A method for accurate recognition of user-defined commands in a ship CAD provided by the first aspect of the present invention includes: Step S100, establishing a hot word table based on professional terms in the ship design field, CAD operation instruction keywords, material specifications, and dimensional parameters; for example, "hatch", "stiffener", etc.

[0041] In step S100, the update method of the hot word table includes:

[0042] When the user corrects the recognition error, the system records the error case and generates a training data set, and uses an algorithm to optimize the hot word weight. The weight adjustment formula is as follows:

[0043]

[0044] where α is the learning rate, R is the actual recognition accuracy rate, is the expected value, Sim is the semantic similarity function, and β is the correction coefficient;

[0045] When the user inputs a non-listed word, the system calculates its semantic similarity with the hot word library through a word vector model, and uses it as a candidate word for the user to confirm and then add it to the hot word table.

[0046] Step S200: Assign weights to each hot word in the hot word list.

[0047] For example, the range of hot word weights is [1 - 5, 10]. The larger the value, the higher the recognition priority of the hot word. Among them, weights 1 - 5 are suitable for general hot words that need to improve the recognition probability but do not affect the recognition of regular words. The weight 10 is used as a super hot word, which is suitable for instructions that may cause significant operation deviations when incorrect (such as key modeling instructions like "create", "connect", "split", etc.), ensuring that it almost certainly appears during recognition. The hot word list is stored in the JSON array format, and the fields include "word" (hot word text) and "weight" (hot word weight).

[0048] Step S300: The user enters a voice command through a collection device; it can be a microphone.

[0049] Step S400: Receive the user's voice command and decode it. When the content included in the hot word list is entered by the user, the decoder, through the weight assignment of the hot word, tends to output professional vocabulary in the shipbuilding industry.

[0050] In the above solution, it should be noted that:

[0051] Adjust the decoder of the speech model, input the hot word list, and dynamically adjust the hot word probability. This solution selects an end-to-end ASR basic model (such as the Transformer ASR class model) as the speech recognition model, and uses the method of combining Beam Search decoding with shallow fusion in the decoding stage to dynamically adjust the probability of each candidate word during decoding. The specific implementation steps are as follows:

[0052] Introduce a hot word enhancement mechanism during the beam search process of speech recognition decoding. Beam Search maintains several candidate sequences of the most probable words, and each word has a corresponding weight score. By adjusting the weight value of the hot word, the score of the word in the candidate sequence can be changed, so that the "hot word weight" is applied to the output result to achieve a method of enhancing the probability of specific words during inference. Specifically, give an additional score bonus to the candidate words that appear in the hot word list to increase their probability in the decoding output and ensure that these specific words win in the competition of the candidate sequences.

[0053] Map the weight value to a fixed probability bonus. For general hot words (weights 1 - 5), set a corresponding +0.1 log probability for each level. For super hot words (weight 10), give a +2.0 log probability boost to ensure that the word is almost certainly output by the model on the premise that it appears in the candidate sequence.

[0054] Technical effects achieved by the above solution in the present invention: Through the above deployment optimization, the entire speech recognition system can run efficiently and stably locally. When the user issues a voice command in the CAD software, the background service quickly transcribes the voice into a text instruction for output, and professional terms are accurately restored due to hot word enhancement. This context bias does not require retraining the model and only acts in the inference stage, so it is very efficient and the hot word list can be adjusted at any time as needed without affecting the basic model. In addition, offline operation not only ensures a low-latency experience but also avoids the risk of uploading sensitive design data to the cloud.

[0055] In one embodiment, for the voice instruction input by the user, an end-to-end ASR model is used for voice-to-text conversion; the present invention also relates to the fusion of multi-modal input signals, including capturing the user's gesture actions, using the OpenPose algorithm to identify the gesture type and align it with the voice instruction timing; obtaining the user's line-of-sight focus area and locking the operation object to the model component within the current window; constructing a multi-modal instruction parsing engine to integrate voice, gesture, and eye movement signals into a structured instruction through a timestamp synchronization mechanism.

[0056] The obtaining of the user's line-of-sight focus area and locking the operation object to the model component within the current window includes the combined binding of the eye movement focus area and the voice instruction object, specifically: if the user mentions ship structure A in the voice instruction and the eye movement focus is in the area of structure A, the operation object is automatically locked to component A; if the eye movement focus area does not match the position of structure A mentioned in the voice instruction, a confirmation mechanism is triggered to provide candidate word corresponding components for the user to select.

[0057] In one embodiment, the operation state of the CAD software is obtained in real time, including the current editing layer, the activated tool type, and the model component attributes, and the hot word weights are dynamically adjusted according to the operation state.

[0058] For example, if the user is in the "pipe layout" mode, the weights of pipe-related terms such as "flange" and "elbow" are increased to the super hot word level; if the user continuously executes the instructions "create rib plate" → "adjust thickness", the associated words of "material grade" and "welding method" are pre-loaded into the buffer.

[0059] The following information is obtained in real time through the CAD software API (such as AutoCAD.NET or SolidWorks API): the currently activated tool (such as "extrusion" "cutting"); the selected component attributes (such as material, size, belonging subsystem); the view parameters (such as viewing angle, zoom ratio); the user's historical operation sequence is analyzed to predict the user's future n-step instruction input, and the relevant hot words are pre-loaded into the buffer. Based on the LSTM model, the user's operation sequence (such as "create → modify → save") is analyzed to predict future instructions and pre-load relevant hot words to shorten the response time.

[0060] For a better understanding of the present invention, the following further elaborates on the present invention through simple examples.

[0061] Embodiment: Take the pipeline layout as an example.

[0062] User voice command: "Add a DN200 flange at the third rib position";

[0063] Gesture action: Circle the target rib position area;

[0064] System action: Identify "rib position", "DN200", and "flange" as super hot words to ensure accurate transcription; Combine the gesture circled area to locate the operation object, and automatically fill in the flange standard parameters according to the current "pipeline design" mode.

[0065] The user inputs a new term "collision bulkhead", and the system recommends related hot words "bulkhead" and "stiffener"; After the user selects "bulkhead", the system adds this word to the hot word library and assigns an initial weight.

[0066] A ship CAD user-defined command precise recognition system proposed in the second aspect of the present invention includes: a hot word table construction module for establishing a hot word table based on professional terms in the ship design field, CAD operation instruction keywords, material specifications, and dimensional parameters; a weight management module for assigning weights to each hot word in the hot word table; a collection module for the user to input voice commands through a collection device; an output module for receiving the user voice commands and decoding them. When the content included in the hot word table is input by the user, the decoder tends to output professional terms in the ship industry through the weight assignment of the hot words.

[0067] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for accurately identifying user-defined commands in ship CAD, characterized in that, Including: Step S100, establishing a hot word list based on professional terms in the ship design field, CAD operation instruction keywords, material specifications, and dimensional parameters; Step S200, assigning weights to each hot word in the hot word list; Step S300, the user enters a voice command through a collection device; Step S400, receiving the user's voice command and decoding it. When the content included in the hot word list entered by the user is recognized, the decoder, through the weight assignment of the hot word, tends to output professional vocabulary in the ship industry.

2. The precise recognition method for user-defined commands in ship CAD according to claim 1, characterized in that The hot word list is stored in the JSON array format. The fields include hot word text and hot word weight. When the voice recognition service is started, the configured JSON hot word list is read, and all hot words and their weights are parsed for the beam search algorithm in the voice recognition decoding to match and adjust the weights.

3. A method for accurately identifying user-defined commands in ship CAD according to claim 2, characterized in that, Introduce a hot word enhancement mechanism during the beam search process of voice recognition decoding to maintain several candidate sequences of the most probable words, where each word has a corresponding weight score. By adjusting the weight value of the hot word, the score of the word in the candidate sequence is changed; By applying the hot word weight to the output result, the probability of specific words is enhanced during inference; By giving an additional score bonus to the candidate words that appear in the hot word list, the probability of the candidate words in the decoding output is increased.

4. A method for accurately identifying user-defined commands in ship CAD according to claim 3, characterized in that Map the weight value to a fixed probability bonus.

5. A method for accurately identifying user-defined commands in ship CAD according to claim 1, characterized in that, For the voice command entered by the user, use an end-to-end ASR model for voice text conversion; also including: Capture the user's gesture actions, use the OpenPose algorithm to identify the gesture type and align it with the voice command timing; Obtain the user's line of sight focus area, and lock the operation object to the model component within the current window; Build a multimodal instruction parsing engine, and integrate voice, gesture, and eye movement signals into a structured instruction through a timestamp synchronization mechanism.

6. A method for accurately identifying user-defined commands in ship CAD according to claim 1, characterized in that Obtain the operation status of the CAD software in real time, including the current editing layer, the activated tool type, and the model component attributes, and dynamically adjust the hot word weights according to the operation status.

7. A method for accurately identifying user-defined commands in ship CAD according to claim 1, characterized in that In step S100, the update method of the hot word list includes: When the user corrects the recognition error, the system records the error case and generates a training data set, and uses an algorithm to optimize the hot word weight. The weight adjustment formula is as follows: Among them, α is the learning rate, R is the actual recognition accuracy rate, is the expected value, Sim is the semantic similarity function, and β is the correction coefficient; When the user enters an uncollected word, the system calculates its semantic similarity with the hot word library through a word vector model, and uses it as a candidate word for the user to confirm and then add it to the hot word list.

8. A method for accurately identifying user-defined commands in ship CAD according to claim 1, characterized in that, It also includes analyzing the user's historical operation sequence, predicting the user's future n-step instruction entry, and preloading relevant hot words into the buffer.

9. A method for accurately identifying user-defined commands in ship CAD according to claim 5, characterized in that The obtaining of the user's line of sight focus area and locking the operation object to the model component within the current window includes the combined binding of the eye movement focus area and the voice command object. Specifically: If the user mentions ship structure A in the voice command and the eye movement focus is in the area of structure A, the operation object is automatically locked as component A; If the eye movement focus area does not match the position of structure A mentioned in the voice command, a confirmation mechanism is triggered, and the component corresponding to the candidate word is provided for the user to select.

10. A precise recognition system for user-defined commands in ship CAD, characterized in that, Including: A hot word list construction module, used to establish a hot word list based on professional terms in the ship design field, CAD operation instruction keywords, material specifications, and dimensional parameters; A weight management module for assigning weights to each hot word in the hot word list; An acquisition module for the user to input voice commands through an acquisition device; An output module for receiving and decoding the user's voice commands. When the content included in the hot word list is input by the user, the decoder tends to output professional vocabulary in the shipping industry through the weight assignment of hot words.