Processing method, terminal device, and storage medium
Patent Information
- Application Number
- CN202610560682.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-24
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]相关技术中,在使用多语种的语言模型时,会加载模型的全部数据,导致使用时对设备内存占用较大
[0015] The processing method, terminal device, and storage medium of this application include: in response to meeting preset conditions, loading target data matching the target language and/or target country from a multilingual language model. The technical solution of this application, when using a multilingual language model, only loads the target data matching the target language and/or target country from the language model, which can reduce memory usage.
Smart Images

Figure CN122594366A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, specifically to a processing method, terminal device, and storage medium. Background Technology
[0002] As the functions of mobile phones, tablets and other terminal devices continue to improve, they have gradually become one of the commonly used tools in people's daily lives and work.
[0003] In related technologies, when using multilingual language models, all the model's data is loaded, resulting in a large memory footprint on the device. Summary of the Invention
[0004] To address the aforementioned technical problems, this application provides a processing method, a terminal device, and a storage medium that can reduce memory usage when using a multilingual language model.
[0005] This application provides a processing method, including: In response to meeting preset conditions, target data matching the target language and / or target country is loaded from the multilingual language model.
[0006] Optionally, load target data from a multilingual language model that matches the target language and / or target country, including: Multiple target lexical units are identified based on the target language and / or target country; Load target data corresponding to multiple target lexical units from a multilingual language model. The target data includes the target embedding layer weight vector corresponding to each target lexical unit.
[0007] Optionally, before determining multiple target lexical units based on the target language and / or target country, the process may also include: Acquire languages and / or corpora from different languages and / or different countries; Segment the language and / or corpus to identify high-frequency word units; Determine the correspondence between word elements and languages and / or countries based on high-frequency word elements and the languages and / or countries they correspond to.
[0008] Optionally, the processing method also includes: The text content of the processing instructions is segmented into words to determine the word combinations corresponding to the text content; Find the target word that corresponds to the word combination among multiple target words, and determine the embedding layer weight vector of the target word that corresponds to the word combination as the embedding layer weight vector of the word combination. The semantic understanding of the text content of the processing instruction is performed based on the embedding layer weight vector corresponding to the word combination, and the processing result of the text content of the processing instruction is output.
[0009] Optionally, preset conditions are met, including: Received a processing instruction in the target language, and / or, located in the target country and received a processing instruction.
[0010] Optionally, the processing method also includes: Processing instructions based on target data.
[0011] Optionally, the processing instructions include: Instructions for processing the input text content; And / or, instructions that process the input target content and output the text content.
[0012] Optionally, the target content includes at least one of images, audio, and video.
[0013] This application also provides a terminal device, including a memory and a processor, wherein the memory stores a processing program or instructions, and when the processing program or instructions are executed by the processor, they implement any of the processing methods described above.
[0014] This application also provides a storage medium storing a computer program or instructions, which, when executed by a terminal device, implement any of the above processing methods.
[0015] The processing method, terminal device, and storage medium of this application include: in response to meeting preset conditions, loading target data matching the target language and / or target country from a multilingual language model. The technical solution of this application, when using a multilingual language model, only loads the target data matching the target language and / or target country from the language model, which can reduce memory usage. Attached Figure Description The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0016] Figure 1 A schematic diagram of the hardware structure of a terminal device for implementing various embodiments of this application.
[0017] Figure 2 This is a communication network system architecture diagram provided for an embodiment of this application.
[0018] Figure 3 This is a flowchart illustrating the processing method according to the first embodiment.
[0019] Figure 4 This is a flowchart illustrating the processing method according to the second embodiment.
[0020] Figure 5 This is a flowchart illustrating the processing method according to the third embodiment.
[0021] The realization of the objectives, functional features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0023] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0024] It should be understood that although the terms first, second, third, etc., may be used herein to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if," as used herein, may be interpreted as "when," "when," or "in response to determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to also include the plural forms unless the context indicates otherwise. It should be further understood that the terms "comprising," "including," indicate the presence of the stated feature, step, operation, element, component, item, kind, and / or group, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms "or," "and / or," "including at least one of the following," etc., as used in this application, may be interpreted as inclusive, or mean any one or any combination thereof. For example, "including at least one of the following: A, B, C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Similarly, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Exceptions to this definition only occur when the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.
[0025] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0026] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrases “if determination” or “if detection (of the stated condition or event)” can be interpreted as “when determination” or “in response to determination” or “when detection (of the stated condition or event)” or “in response to detection (of the stated condition or event).”
[0027] It should be noted that step designations such as S1 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. Those skilled in the art may change the execution order of the steps in specific implementation, but these should all be within the protection scope of this application.
[0028] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0029] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0030] Terminal devices can be implemented in various forms. For example, the terminal devices described in this application may include mobile terminals such as mobile phones, tablets, laptops, PDAs, smartwatches, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.
[0031] Please see Figure 1 This is a schematic diagram of the hardware structure of a terminal device implementing various embodiments of this application. The terminal device 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111, etc. Those skilled in the art will understand that... Figure 1 The terminal device structure shown does not constitute a limitation on the terminal device. The terminal device may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0032] The following is combined Figure 1 A detailed introduction to each component of the terminal device: The radio frequency unit 101 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 110; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, and a duplexer. Furthermore, the radio frequency unit 101 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), 5G, and 6G.
[0033] WiFi is a short-range wireless transmission technology. Terminal devices using the WiFi module 102 can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 1 WiFi module 102 is shown, but it is understood that it is not a necessary component of the terminal device and can be omitted as needed without changing the essence of the invention.
[0034] The audio output unit 103 can convert audio data received by the radio frequency unit 101 or the WiFi module 102 or stored in the memory 109 into audio signals and output them as sound when the terminal device 100 is in call signal receiving mode, call mode, recording mode, voice recognition mode, broadcast receiving mode, etc. Furthermore, the audio output unit 103 can also provide audio output related to specific functions performed by the terminal device 100 (e.g., call signal receiving sound, message receiving sound, etc.). The audio output unit 103 may include a speaker, a buzzer, etc.
[0035] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on the display unit 106. The image frames processed by the GPU 1041 can be stored in the memory 109 (or other storage media) or transmitted via the radio frequency unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) in operating modes such as telephone call mode, recording mode, and voice recognition mode, and can process such sound into audio data. The processed audio (voice) data can be converted into a format that can be transmitted to a mobile communication base station via the radio frequency unit 101 in telephone call mode. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.
[0036] The terminal device 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the ambient light level, and the proximity sensor can turn off the display panel 1061 and / or backlight when the terminal device 100 is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that can also be configured in the phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0037] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0038] User input unit 107 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function control of the terminal device. Optionally, user input unit 107 may include touch panel 1071 and other input devices 1072. Touch panel 1071, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 1071), and drive corresponding connection devices according to a pre-set program. Touch panel 1071 may include two parts: a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to processor 110, and can receive and execute commands sent by processor 110. In addition, touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may also include other input devices 1072. Optionally, other input devices 1072 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc., without being specifically limited here.
[0039] Optionally, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. Subsequently, the processor 110 provides corresponding visual output on the display panel 1061 based on the type of touch event. Although in Figure 1 In this embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the terminal device. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the terminal device. The specific implementation is not limited here.
[0040] Interface unit 108 serves as an interface through which at least one external device can connect to terminal device 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 108 may be used to receive input (e.g., data, power, etc.) from the external device and transmit the received input to one or more elements within terminal device 100, or it may be used to transmit data between terminal device 100 and the external device.
[0041] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a program storage area and a data storage area. Optionally, the program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 109 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0042] The processor 110 is the control center of the terminal device. It connects various parts of the terminal device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 109, and by calling data stored in the memory 109, it performs various functions and processes data of the terminal device, thereby providing overall monitoring of the terminal device. The processor 110 may include one or more processing units; preferably, the processor 110 may integrate an application processor and a modem processor. Optionally, the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 110.
[0043] The terminal device 100 may also include a power supply 111 (such as a battery) that supplies power to various components. Preferably, the power supply 111 can be logically connected to the processor 110 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0044] although Figure 1 As not shown, the terminal device 100 may also include a Bluetooth module, etc., which will not be described in detail here.
[0045] To facilitate understanding of the embodiments of this application, the communication network system on which the terminal device of this application is based is described below.
[0046] Please see Figure 2 , Figure 2 This application provides a communication network system architecture diagram. The communication network system is an LTE system based on the universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203, and the operator's IP services 204, which are connected in sequence.
[0047] Optionally, UE201 can be the aforementioned terminal 100, which will not be described in detail here.
[0048] E-UTRAN202 includes eNodeB2021 and other eNodeB2022s. Optionally, eNodeB2021 can connect to other eNodeB2022s via backhaul (e.g., X2 interface). eNodeB2021 connects to EPC203 and can provide UE201 with access to EPC203.
[0049] EPC203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving Gateway) 2034, a PGW (Packet Data Network Gateway) 2035, and a PCRF (Policy and Charging Rules Function) 2036, etc. Optionally, MME2031 is the control node that handles signaling between UE201 and EPC203, providing bearer and connection management. HSS2032 is used to provide registers to manage functions such as the Home Location Register (not shown in the figure) and stores user-specific information such as service characteristics and data rates. All user data can be sent through SGW2034. PGW2035 can provide UE 201 IP address allocation and other functions. PCRF2036 is the policy and charging control decision point for service data flow and IP bearer resources. It selects and provides available policy and charging control decisions for the policy and charging enforcement function unit (not shown in the figure).
[0050] IP services 204 may include the Internet, intranet, IMS (IP Multimedia Subsystem), or other IP services.
[0051] Although the above description uses the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but also to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., without limitation.
[0052] Based on the aforementioned terminal device hardware structure and communication network system, various embodiments of this application are proposed.
[0053] First Embodiment Reference Figure 3 , Figure 3 This is a flowchart illustrating the processing method according to the first embodiment. The processing method of this application embodiment can be applied to terminal devices (such as mobile phones) and includes the following steps: S1: In response to meeting preset conditions, load target data from the multilingual language model that matches the target language and / or target country.
[0054] Optionally, when preset conditions are met, a data loading instruction is triggered to begin loading target data. Optionally, the target data is loaded into the memory of the terminal device so that it can be used to perform corresponding data processing. For example, in response to the satisfaction of preset conditions, target data matching the target language and / or target country from a multilingual language model is loaded from secondary storage into memory for the processor to read.
[0055] A multilingual language model refers to a large model that supports multilingual language processing. Optionally, a multilingual language model can be an existing open-source large model, or an optimized version of an existing open-source large model. By using a multilingual language model, a wider range of languages can be applied, and it is not necessary to train a separate model for each country or language, thus reducing training costs and lowering overall costs.
[0056] Optionally, target data can be determined by matching the target language, or by matching the target country, or by matching both the target language and the target country.
[0057] Optionally, when a country uses a relatively uniform language, the target data can be determined by matching the target language or the target country. In this way, only a simpler correspondence needs to be established between the target data and the language or country, reducing the overall data size of the language model.
[0058] Optionally, when a country uses a variety of languages, the target data can be determined by matching the target language with the target country. Optionally, when determining target data by matching the target language with the target country, first determine the data corresponding to the target country, and then match the data corresponding to the target language from the data corresponding to the target country, thus obtaining the target data. This method reduces the range of data to be searched for the target language, and then loads data specifically for the target language within the data range, accelerating the data loading speed.
[0059] The technical solution of this application reduces memory usage by loading only the target data that matches the target language and / or target country in the language model when using a multilingual language model.
[0060] Optionally, preset conditions are met, including: Received a processing instruction in the target language, and / or, located in the target country and received a processing instruction.
[0061] Optionally, using a processing instruction in the target language means that the processing instruction is input in the target language. The target language is the language used in the processing instruction. For example, if the processing instruction is "translate", then the target language is Chinese. Optionally, the processing instruction can be a voice instruction input via voice input or a text instruction input through an input box.
[0062] Optionally, the current country of the terminal device can be determined through its location, which is the target country. When a processing instruction is received, the target country can be determined by obtaining the current location. Optionally, the target data can be determined directly by matching the target country, or the target language can be determined based on the target country, and then the target data can be determined by matching the target language.
[0063] Optionally, in step S1, target data matching the target language and / or target country is loaded from the multilingual language model, including: Multiple target lexical units are identified based on the target language and / or target country; Load target data corresponding to multiple target lexical units from a multilingual language model. The target data includes the target embedding layer weight vector corresponding to each target lexical unit.
[0064] Optionally, a pre-established correspondence is established between the tokens of the language model and the language and / or country, and each token has a corresponding embedding weight vector. Tokens are the basic semantic units (such as words, subwords, or characters) segmented from text, and the embedding weight vectors are learnable parameters used in the model to map these tokens into high-dimensional vectors. Together, they constitute the semantic representation foundation of the language model.
[0065] Optionally, based on the correspondence between lexical units and languages and / or countries, multiple lexical units corresponding to the target language and / or target country can be determined, and these multiple lexical units are the multiple target lexical units. After determining the multiple target lexical units, the data in the language model corresponding to the multiple target lexical units is the target data. Optionally, the target data includes at least the embedding layer weight vector corresponding to each target lexical unit, and these embedding layer weight vectors are the target embedding layer weight vectors.
[0066] By loading only the target embedding layer weight vectors corresponding to the target language and / or target country, compared to loading the entire model's data, the memory usage of data loading can be effectively reduced, and the loading speed can be improved. This is suitable for mobile devices, edge devices, and low-bandwidth scenarios.
[0067] Taking the sentence-transformers / paraphrase-multilingual-MiniLM-L12-v2 language model as an example, this model has 96 million parameters in the embedding layer and 21 million parameters in the transformer network layer, including 250,000 tokens, and supports more than 10 languages. When loading data for only one language, the data loading ratio is (9.6 million * 4 + 21 million * 4) / (96 million * 4 + 21 million * 4) = 0.26, which can save 74% of memory usage.
[0068] Optionally, before determining multiple target lexical units based on the target language and / or target country to pre-establish the correspondence between lexical units and languages and / or countries, the method further includes: Acquire languages and / or corpora from different languages and / or different countries; Segment the language and / or corpus to identify high-frequency word units; Determine the correspondence between word elements and languages and / or countries based on high-frequency word elements and the languages and / or countries they correspond to.
[0069] Optionally, the languages and / or corpora can be obtained from public corpora, user behavior logs, social media, or localized text (such as official documents, e-commerce product descriptions, etc.). Optionally, the languages and / or corpora of different languages are each for a specific language. Optionally, the languages and / or corpora of different countries may include at least one language, depending on the number of languages used in that country.
[0070] Optionally, when segmenting languages and / or corpora, an appropriate segmenter can be selected based on the language to improve the accuracy of the segmentation results. Alternatively, a unified multilingual segmenter can be used, eliminating reliance on a specific language and broadening the range of languages that can be used.
[0071] Optionally, high-frequency words are those that appear more than a preset number of times, indicating words that appear relatively frequently in that language or country. By retaining only high-frequency words, computational overhead and the amount of data related to the constructed correspondences in the language model can be reduced while covering most usage scenarios.
[0072] Optionally, after determining high-frequency word units based on the language and / or corpus of different languages, the high-frequency word units can be used as word units of the corresponding language to establish a correspondence between them. Optionally, the correspondence between word units and languages can be stored by creating language tags for the high-frequency word units.
[0073] Optionally, after determining high-frequency word units based on the languages and / or corpora of different countries, the high-frequency word units can be used as word units of the corresponding countries to establish a correspondence. Optionally, the correspondence between word units and countries can be stored by creating country tags for the high-frequency word units.
[0074] Optionally, the languages and / or corpora of different countries may include at least one language and / or corpus. Therefore, a correspondence between word units and corresponding countries and languages can also be established at the same time.
[0075] By establishing pre-defined correspondences between lexical units and languages and / or countries based on a multilingual language model, target data can be quickly located and loaded when needed, particularly in the target language and / or country. Furthermore, dynamic updates are better supported; when new corpus is available, the correspondences can be re-established without retraining the entire model.
[0076] Second Embodiment Reference Figure 4 , Figure 4 The flowchart illustrating the processing method according to the second embodiment is shown below. The processing method of this application embodiment further includes the following steps: S10, segment the text content of the processing instruction and determine the word combination corresponding to the text content; Optionally, the text content of the processing instruction refers to the content of the input processing instruction. For example, if the user inputs "family" in the function interface, the text content of the processing instruction will be "family".
[0077] By segmenting the text content of the processing instructions and determining the corresponding word combinations, unstructured natural language text can be converted into a discrete index sequence that the model can process. Optionally, the text content can be segmented first according to a preset word segmentation algorithm. For languages without spaces, such as Chinese, word boundaries need to be identified using a statistical or deep learning-based word segmenter; for languages such as English, sub-words are directly split. Next, each segmented sub-word or character is mapped to a unique Token ID in the vocabulary. Finally, an ordered sequence of Token IDs is generated, which represents the word combinations corresponding to the text content and serves as the key for subsequently searching the embedding layer weight vector.
[0078] S20, Search for the target word that corresponds to the word combination among multiple target word words, and determine the embedding layer weight vector of the target word that corresponds to the word combination as the embedding layer weight vector corresponding to the word combination. Optionally, the processing instructions specify the target language and / or target country, and multiple target lexical units are lexical units corresponding to the target language and / or target country. Among these lexical units, the target lexical unit corresponding to the lexical unit combination is searched.
[0079] Optionally, the token ID sequence generated in step S10 can be traversed to determine whether each ID belongs to the set of "multiple target words" of the current target language or country. If a match is found, the embedding layer weight vector of the word can be extracted from the target data. After the traversal is completed, the embedding layer weight vector corresponding to the word combination can be determined.
[0080] S30: Based on the embedding layer weight vector corresponding to the word combination, perform semantic understanding on the text content of the processing instruction, and output the processing result of the text content of the processing instruction.
[0081] Optionally, the attention mechanism of the Transformer architecture can be used to perform deep semantic computation based on the assembled embedding layer weight vectors and output the processing results. This processing procedure is known to those skilled in the art and will not be described in detail here.
[0082] Taking text-based image search as an example, when a user enters "family" in the function interface, the system recognizes that the language of the keyword "family" is English. It then searches the language model for the embedding layer weight vectors corresponding to the English tag words, completes semantic understanding, selects images related to "family", and displays them to the user.
[0083] The above approach reduces the amount of data required to process instructions, decreases memory usage, speeds up the output of processing results, and improves the user experience.
[0084] Third Embodiment Reference Figure 5 , Figure 5 This is a flowchart illustrating the processing method according to the third embodiment. The processing method of this application embodiment further includes the following steps: S100 processes the processing instructions based on the target data.
[0085] Optionally, the target data is data from a multilingual language model that matches the target language and / or target country. Optionally, the processing instruction is an instruction input using the target language, or it can be an instruction received in the target country. Therefore, the processing instruction can specify the target language and / or target country. Processing instructions based on the target data eliminates the need to load all data from the multilingual language model, reducing memory usage, improving response speed, and thus enhancing the user experience.
[0086] Optionally, the processing instructions include: instructions to process the input text content; and / or, instructions to process the input target content and then output the text content.
[0087] Optionally, when the processing instruction is an instruction to process the input text content, the text content is entered together with the processing instruction. The text content can be the task content indicated by the instruction, such as "Find me some pictures of flowers", or the content to be processed by the instruction, such as "Translate + the sentence to be translated".
[0088] Optionally, when the processing instruction is an instruction to process the input target content and output text content, the target content can be input and the processing instruction generated under a specific processing function. Optionally, the specific processing function refers to the application function of the language model, such as audio to document generation, video to document generation, image to document generation, etc. Optionally, the target content includes at least one of images, audio, and video.
[0089] Optionally, under a specific processing function, the target language indicated by the processing instruction can be determined based on the target content, or the target language can be directly selected under a specific processing function. Optionally, when images, audio, or video are input as target content, if the images, audio, or video contain information that can indicate the language, the target language can be confirmed based on the target content. For example, if Chinese audio is input, the target language is Chinese.
[0090] In this way, the target data can be used to process different types of processing instructions, reducing memory usage, improving response speed, and expanding the range of application scenarios.
[0091] Optionally, the processing instructions are processed based on the target data, including: The text content of the processing instructions is segmented to determine the word combinations of the text content; Determine the embedding layer weight vector corresponding to word combination in the target data; Semantic understanding is performed based on the embedding layer weight vector, and the processing results of the output processing instructions are generated.
[0092] Optionally, the text content of the processing instruction refers to the content of the input processing instruction. For example, if the user inputs "family" in the function interface, the text content of the processing instruction will be "family".
[0093] Optionally, by segmenting the text content of the processing instructions to determine the word combinations corresponding to the text content, unstructured natural language text can be converted into a discrete index sequence that the model can process. Optionally, the text content can be segmented first according to a preset word segmentation algorithm. Then, in the target data, target words that match the word combinations are found, and the embedding layer weight vectors of these target words are used as the embedding layer weight vectors corresponding to the word combinations. Finally, the attention mechanism of the Transformer architecture can be used to perform deep semantic computation based on the embedding layer weight vectors corresponding to the word combinations, and the processing result can be output.
[0094] The processing method, terminal device, and storage medium of this application include: in response to meeting preset conditions, loading target data matching the target language and / or target country from a multilingual language model. The technical solution of this application, when using a multilingual language model, only loads the target data matching the target language and / or target country from the language model, which can reduce memory usage, improve response speed, and enhance user experience.
[0095] This application also provides a terminal device, including a memory and a processor, wherein the memory stores a processing program or instructions, and the processing program or instructions, when executed by the processor, implement the method described in any of the above embodiments.
[0096] This application also provides a storage medium storing a computer program or instructions, which, when executed by a terminal device, implements the method described in any of the above embodiments.
[0097] In the embodiments of the terminal device and storage medium provided in this application, all the technical features of any of the above-described processing method embodiments may be included. The extended and explanatory content of the specification is basically the same as that of the embodiments of the above methods, and will not be repeated here.
[0098] This application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to perform the methods described in the various possible implementations above.
[0099] This application also provides a chip, including a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that a device with the chip installed performs the methods described in the various possible implementations above.
[0100] It is understood that the above scenarios are merely examples and do not constitute a limitation on the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, as those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0101] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0102] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.
[0103] The units in the device of this application embodiment can be merged, divided, and deleted according to actual needs.
[0104] In this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions are generally described in detail only when they appear for the first time. When they appear again, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions that are not described in detail later can be referred to their previous relevant detailed descriptions.
[0105] In this application, the descriptions of the various embodiments have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0106] The technical features of the present application can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present application.
[0107] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of this application.
[0108] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, storage disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0109] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A processing method, characterized in that, Including the following steps: In response to meeting preset conditions, target data matching the target language and / or target country is loaded from the multilingual language model.
2. The method according to claim 1, characterized in that, The target data in the multilingual language model that matches the target language and / or target country includes: Multiple target lexical units are determined based on the target language and / or the target country; Load the target data corresponding to the multiple target lexical units in the multilingual language model, wherein the target data includes the target embedding layer weight vector corresponding to each target lexical unit.
3. The method according to claim 2, characterized in that, Before determining multiple target lexical units based on the target language and / or the target country, the method further includes: Acquire languages and / or corpora from different languages and / or countries; segment the languages and / or corpora into words and determine high-frequency word units; Based on the high-frequency lexical units and the corresponding languages and / or countries, determine the correspondence between the lexical units and the languages and / or countries.
4. The method according to claim 3, characterized in that, The method further includes: The text content of the processing instruction is segmented into words to determine the word combination corresponding to the text content; Find the target word corresponding to the word combination among the plurality of target words, and determine the embedding layer weight vector of the corresponding target word as the embedding layer weight vector corresponding to the word combination; Based on the embedding layer weight vector corresponding to the word combination, the semantic understanding of the text content of the processing instruction is performed, and the processing result of the text content of the processing instruction is output.
5. The method according to any one of claims 1 to 4, characterized in that, The conditions for satisfying the preset conditions include: The system receives a processing instruction using the target language, and / or is located in the target country and receives a processing instruction.
6. The method according to any one of claims 1 to 4, characterized in that, The method further includes: The processing instructions are processed based on the target data.
7. The method according to claim 6, characterized in that, The processing instructions include: Instructions for processing the input text content; And / or, instructions that process the input target content and output the text content.
8. The method according to claim 7, characterized in that, The target content includes at least one of images, audio, and video.
9. A terminal device, characterized in that, It includes a memory and a processor, the memory storing a program or instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium stores a computer program or instructions, which, when executed by a terminal device, implement the method as described in any one of claims 1 to 8.