Abstract
Automatic recognition of spoken alphabets and digits is one of the difficult tasks in the field of computer speech recognition. Spoken alphadigits (i.e., alphabets and digits) recognition process is needed in many applications that take spoken digits and/or alphabets as inputs. Arabic language is a Semitic language that differs from other languages such as English. One of these differences is how to pronounce the ten digits and all alphabets. In this research, spoken Arabic digits are investigated from the speech recognition point of view. The system was designed to recognize an isolated whole-word speech based on Hidden Morkov Models (HMM). The designed HMM model was based on phoneme recognition. In the training and testing phase of the system, the Arabic speech corpus known as Saudi Accented Arabic Voice Bank (SAAVB) was used. Nine different experiments were performed on SAAVB database in this research. The first three were trained and tested by using each individual digital subset. The fourth one was conducted on these three subsets collectively (i.e., trained by using all three training subsets and tested by using all three testing subsets). In the following three experiments, the training subset was the same as that of the fourth experiment but the testing subsets were the same as that of the first three experiments. The eighth experiment was on the Arabic alphabets, and the ninth one was applied on the digits and the alphabets collectively.
The research has three main phases; first designing the system by using only the Arabic digits, second working on the Arabic alphabets to be recognized, analyzed, and evaluated, third the combination of the Arabic digits and alphabets together. The system achieved 94.13% overall correct digit recognition in the noisy environment using mixed training and testing subsets collectively. In the case of the alphabet subsets, the overall system performance was 64.06%, which is reasonably high where our database consists of a noisy corpus. With mixed alphabets and digits, the overall system accuracy was 76.06% which is better than the alphabet experiments but much less than those conducted for the digits. The paper includes some future work suggestions in addition to recommendations for the SAAVB design and management community at KACST.