Muft Shiksha™ एक 100% Free Education Portal है 🇮🇳, जिसका उद्देश्य Class 9–12 के हर विद्यार्थी तक High-Quality Education को पूरी तरह मुफ्त पहुँचाना है। 🇮🇳 हम मानते हैं कि अच्छी शिक्षा किसी student की आर्थिक स्थिति पर निर्भर नहीं होनी चाहिए। 🇮🇳 हर विद्यार्थी को वही Quality Study Material, MCQs, Quizzes, Exam Preparation, Concept-Based Learning और Bilingual Support मिलना चाहिए, जो आमतौर पर महंगी Coaching या Premium Platforms में मिलता है। Muft Shiksha™ 🇮🇳 इसी सोच के साथ बनाया गया है • Muft Shiksha™ एक 100% Free Education Portal है 🇮🇳, जिसका उद्देश्य Class 9–12 के हर विद्यार्थी तक High-Quality Education को पूरी तरह मुफ्त पहुँचाना है। 🇮🇳 हम मानते हैं कि अच्छी शिक्षा किसी student की आर्थिक स्थिति पर निर्भर नहीं होनी चाहिए। 🇮🇳 हर विद्यार्थी को वही Quality Study Material, MCQs, Quizzes, Exam Preparation, Concept-Based Learning और Bilingual Support मिलना चाहिए, जो आमतौर पर महंगी Coaching या Premium Platforms में मिलता है। Muft Shiksha™ 🇮🇳 इसी सोच के साथ बनाया गया है • Muft Shiksha™ एक 100% Free Education Portal है 🇮🇳, जिसका उद्देश्य Class 9–12 के हर विद्यार्थी तक High-Quality Education को पूरी तरह मुफ्त पहुँचाना है। 🇮🇳 हम मानते हैं कि अच्छी शिक्षा किसी student की आर्थिक स्थिति पर निर्भर नहीं होनी चाहिए। 🇮🇳 हर विद्यार्थी को वही Quality Study Material, MCQs, Quizzes, Exam Preparation, Concept-Based Learning और Bilingual Support मिलना चाहिए, जो आमतौर पर महंगी Coaching या Premium Platforms में मिलता है। Muft Shiksha™ 🇮🇳 इसी सोच के साथ बनाया गया है
data science methodology MCQ Questions for Class 12
data science methodology se related questions ko ek jagah revise karein. Har question me bilingual content, answer feedback aur explanation available hai.
A. प्रतिरूप बहुसंख्यक वर्ग के पक्ष में झुक सकता है/The model may become biased toward the majority class
Explanation
Simple Explanation
असंतुलित आँकड़ों में एक वर्ग के उदाहरण अन्य वर्गों की तुलना में बहुत अधिक होते हैं। इसलिए प्रतिरूप बहुसंख्यक वर्ग के उदाहरणों को अधिक सीखकर उसके पक्ष में झुक सकता है और अल्पसंख्यक वर्ग को कम सही पहचान सकता है। विकल्प B गलत है, क्योंकि आँकड़ों का असंतुलन अपने-आप सभी वर्गों के लिए निष्पक्ष प्रदर्शन सुनिश्चित नहीं करता। / In imbalanced data, one class has many more examples than the other class or classes. The model may therefore learn the majority class more strongly, become biased toward it, and identify the minority class less accurately. Option B is incorrect because class imbalance does not automatically ensure fair performance for every class.
A. स्रोत प्रामाणिक हो, जानकारी सत्यापित हो और वह अद्यतन हो/The source is authoritative, its information is verified, and it is up to date
Explanation
Simple Explanation
विश्वसनीय डेटा स्रोत की प्रामाणिकता जानी जा सकती है, उसकी जानकारी सत्यापित होती है और वह वर्तमान जानकारी के अनुसार अद्यतन रहता है। इसलिए विकल्प A सही है। इसके विपरीत, लेखक या संस्था का विवरण तथा प्रकाशन/अद्यतन तिथि न होना स्रोत की जवाबदेही और अद्यतनता का आकलन कठिन बनाता है। / A reliable data source has identifiable authority, contains verified information, and is current or regularly updated. Therefore, option A is correct. In contrast, the absence of author or organization details and publication or update dates makes it difficult to assess a source's accountability and currency.
A. क्या समाधान परियोजना के निर्धारित उद्देश्यों और अपेक्षित परिणामों को प्राप्त कर रहा है/Whether the solution is achieving the project’s defined objectives and expected outcomes
Explanation
Simple Explanation
सफलता संकेतक पहले से निर्धारित, मापने योग्य मानदंड होते हैं जिनसे यह जाँचा जाता है कि समाधान परियोजना के उद्देश्यों और अपेक्षित परिणामों को प्राप्त कर रहा है या नहीं। एकत्र किए गए डेटा की मात्रा या टीम द्वारा लगाया गया समय परियोजना के आँकड़े हो सकते हैं, लेकिन वे स्वयं सफलता का मुख्य माप नहीं हैं, जब तक उन्हें विशेष रूप से परियोजना के उद्देश्य के रूप में निर्धारित न किया गया हो। / Success indicators are predefined, measurable criteria used to determine whether a solution is achieving the project’s objectives and expected outcomes. The amount of data collected or the time spent by a team may be project statistics, but they are not by themselves the primary measure of success unless they have specifically been set as project objectives.
A. आँकड़ों की संरचना, संबंधों और संभावित समस्याओं को समझना/To understand the structure, relationships, and potential problems in the data
Explanation
Simple Explanation
आँकड़ा अन्वेषण में आँकड़ों की संरचना, प्रकार, वितरण, संबंधों तथा संभावित त्रुटियों या असामान्य मानों का प्रारंभिक अध्ययन किया जाता है। इसलिए विकल्प A सही है। विकल्प B के विपरीत, आँकड़ों की जाँच किए बिना उनका उपयोग करने से गलत निष्कर्ष निकल सकते हैं। / Data exploration is the preliminary examination of a dataset’s structure, types, distributions, relationships, and possible errors or unusual values. Therefore, option A is correct. In contrast to option B, using data without examining it can lead to incorrect conclusions.
A. ताकि वह लक्ष्य जनसंख्या की विशेषताओं को सही ढंग से दर्शा सके/So that it can accurately reflect the characteristics of the target population
Explanation
Simple Explanation
प्रतिनिधि नमूना लक्ष्य जनसंख्या की महत्वपूर्ण विशेषताओं, जैसे आयु, क्षेत्र या अन्य प्रासंगिक समूहों, को उचित रूप से दर्शाता है। इसलिए नमूने से निकाले गए निष्कर्षों को पूरी जनसंख्या पर अधिक विश्वसनीय रूप से लागू किया जा सकता है। विकल्प C गलत है क्योंकि प्रतिनिधि नमूनाकरण पक्षपात को कम कर सकता है, लेकिन हर प्रकार के पक्षपात को पूरी तरह समाप्त नहीं करता। / A representative sample appropriately reflects important characteristics of the target population, such as age, region, or other relevant groups. Therefore, conclusions drawn from the sample can be applied to the whole population more reliably. Option C is incorrect because representative sampling can reduce bias, but it cannot completely eliminate every type of bias.
A. उपयुक्त स्रोतों से आवश्यक डेटा एकत्र करना/Collecting required data from suitable sources
Explanation
Simple Explanation
डेटा अधिग्रहण का अर्थ किसी समस्या या परियोजना के लिए उपयुक्त स्रोतों से आवश्यक डेटा प्राप्त करना या एकत्र करना है। इसलिए विकल्प A सही है। विकल्प B डेटा सफाई से, विकल्प C डेटा विश्लेषण से, और विकल्प D मॉडल प्रशिक्षण से संबंधित है; ये डेटा अधिग्रहण के बाद के अलग-अलग चरण हो सकते हैं। / Data acquisition means obtaining or collecting the data needed for a problem or project from suitable sources. Therefore, option A is correct. Option B refers to data cleaning, option C to data analysis, and option D to model training; these are distinct stages that may occur after data acquisition.
A. ताकि समाधान की सफलता का वस्तुनिष्ठ मूल्यांकन किया जा सके/So that the success of the solution can be evaluated objectively
Explanation
Simple Explanation
मापने योग्य लक्ष्य सफलता का स्पष्ट और वस्तुनिष्ठ मानदंड देता है। समाधान लागू करने के बाद वास्तविक परिणामों की तुलना लक्ष्य से करके यह आंका जा सकता है कि समस्या कितनी प्रभावी ढंग से हल हुई। विकल्प B के विपरीत, स्पष्ट दायरा और मापदंड समस्या को समझने तथा उसका मूल्यांकन करने में सहायता करते हैं। / A measurable goal provides a clear, objective criterion for success. After implementing a solution, the actual results can be compared with the target to assess how effectively the problem has been solved. Unlike option B, a clear scope and measurable criteria help in understanding and evaluating the problem.
B. समस्या के उद्देश्य, अपेक्षित परिणाम और दायरे को स्पष्ट करना/To clarify the problem goal, expected outcome, and scope
Explanation
Simple Explanation
समस्या-निर्धारण में यह स्पष्ट किया जाता है कि किस समस्या को हल करना है, अपेक्षित परिणाम क्या है और परियोजना की सीमाएँ क्या हैं। इसलिए विकल्प B सही है। विकल्प C में डेटा एकत्र करना समस्या की आवश्यकताओं से निर्देशित होना चाहिए; केवल अधिक डेटा एकत्र करना समस्या-निर्धारण का उद्देश्य नहीं है। / Problem scoping clarifies the problem to be solved, the expected outcome, and the boundaries of the project. Therefore, option B is correct. In option C, data collection should be guided by the problem requirements; simply collecting more data is not the purpose of problem scoping.
A. सीखने और सुधार के अवसरों की पहचान करने के लिए/To identify opportunities for learning and improvement
Explanation
Simple Explanation
अंतिम समीक्षा में परियोजना के लक्ष्यों, परिणामों, सीमाओं, त्रुटियों और प्राप्त प्रतिक्रिया का मूल्यांकन किया जाता है। इससे यह पता चलता है कि क्या अच्छा हुआ और भविष्य में किन बातों में सुधार किया जा सकता है। इसलिए विकल्प A सही है। विकल्प B के विपरीत, समीक्षा में परिणामों की जाँच की जाती है; विकल्प D के विपरीत, त्रुटियों की पहचान करके उन्हें सुधारने के अवसर खोजे जाते हैं। / A final review evaluates the project’s goals, results, limitations, errors, and feedback. It shows what worked well and what can be improved in future work. Therefore, option A is correct. Unlike option B, a review checks the results; unlike option D, it identifies errors and opportunities to address them.
A. नए डेटा पर मॉडल की क्षमता का निष्पक्ष आकलन करने के लिए/To evaluate the model’s ability fairly on new data
Explanation
Simple Explanation
मॉडल प्रशिक्षण डेटा से पैटर्न सीखता है, जबकि परीक्षण डेटा प्रशिक्षण के दौरान अनदेखा रहना चाहिए। दोनों को अलग रखने से नए डेटा पर मॉडल की सामान्यीकरण क्षमता और वास्तविक प्रदर्शन का निष्पक्ष आकलन किया जा सकता है। विकल्प B के विपरीत, परीक्षण डेटा को अलग रखने का उद्देश्य प्रशिक्षण रोकना नहीं, बल्कि प्रशिक्षित मॉडल का मूल्यांकन करना है। / A model learns patterns from training data, whereas testing data should remain unseen during training. Keeping them separate enables a fair evaluation of the model’s generalization and actual performance on new data. Unlike option B, the purpose of a separate test set is not to stop training, but to evaluate the trained model.
A. विश्वसनीय और प्रासंगिक आँकड़े प्राप्त करने के लिए/To obtain reliable and relevant data
Explanation
Simple Explanation
स्रोतों का पहले चयन करने से यह सुनिश्चित होता है कि एकत्र किए गए आँकड़े उद्देश्य के लिए प्रासंगिक, विश्वसनीय और उचित गुणवत्ता के हों। अविश्वसनीय या असंगत स्रोत पक्षपाती अथवा गलत निष्कर्ष दे सकते हैं। विकल्प B गलत है, क्योंकि आँकड़ों की अनावश्यक मात्रा से अधिक उनकी प्रासंगिकता और गुणवत्ता महत्वपूर्ण होती है। / Selecting sources in advance helps ensure that the collected data is relevant to the purpose, reliable, and of appropriate quality. Unreliable or unsuitable sources can lead to biased or inaccurate conclusions. Option B is incorrect because data relevance and quality are more important than collecting an unnecessary quantity of data.
A. समस्या-निर्धारण और परियोजना-योजना के प्रारम्भिक चरण में, प्रभावित लोगों की जरूरतें समझने के लिए/During the early problem-scoping and project-planning stage, to understand affected people’s needs
Explanation
Simple Explanation
हितधारक वे व्यक्ति या समूह हैं जो समस्या, समाधान या उसके परिणामों से प्रभावित होते हैं। उनकी पहचान परियोजना के प्रारम्भिक समस्या-निर्धारण और योजना चरण में करने से उपयोगकर्ताओं की जरूरतें, अपेक्षाएँ और संभावित प्रभाव समझे जाते हैं। इसलिए विकल्प A सही है। डेटा एकत्र करने या अंतिम परीक्षण तक प्रतीक्षा करने से जरूरी आवश्यकताएँ छूट सकती हैं। परीक्षा युक्ति: हितधारक पहचान को समस्या की समझ और उपयोगकर्ता-केंद्रित योजना से जोड़ें। / Stakeholders are individuals or groups affected by the problem, the solution, or its outcomes. Identifying them during early problem scoping and planning helps the team understand users’ needs, expectations, and possible impacts. Therefore, option A is correct. Waiting until data collection or final testing may cause important requirements to be missed. Exam tip: Associate stakeholder identification with understanding the problem and planning around users’ needs.
A. स्पष्ट, संक्षिप्त और मापनीय/Clear, concise, and measurable
Explanation
Simple Explanation
एक प्रभावी समस्या कथन समस्या को स्पष्ट और संक्षिप्त रूप से बताता है तथा सफलता को मापने के लिए मापनीय लक्ष्य देता है। विकल्प B में स्पष्टता नहीं है, जबकि विकल्प C में समस्या के उद्देश्य और लक्षित उपयोगकर्ता का अभाव है। / An effective problem statement describes the problem clearly and concisely and includes measurable goals for judging success. Option B lacks clarity, while option C does not identify the purpose or intended users of the solution.
सही उत्तर स्रोत की विश्वसनीयता है। विश्वसनीय स्रोत के डेटा के सटीक, सत्यापित और भरोसेमंद होने की संभावना अधिक होती है। डेटा की प्रासंगिकता यह बताती है कि डेटा समस्या के लिए उपयोगी है या नहीं, लेकिन वह स्रोत की सत्यता या भरोसेमंदता को सीधे नहीं बताती। / The correct answer is reliability of the source. Data from a reliable source are more likely to be accurate, verifiable, and trustworthy. Relevance indicates whether data are useful for the problem, but it does not directly establish whether the source is trustworthy.
A. सूचित सहमति लेना और गोपनीयता की रक्षा करना/Obtaining informed consent and protecting privacy
Explanation
Simple Explanation
नैतिक आंकड़ा संग्रहण में संबंधित व्यक्ति को स्पष्ट जानकारी देकर उसकी सूचित सहमति लेना तथा उसकी निजी जानकारी की गोपनीयता सुरक्षित रखना आवश्यक है। इसलिए विकल्प A सही है। विकल्प B में अनुमति और पारदर्शिता का अभाव है, जबकि विकल्प C डेटा के उपयोग के बारे में जानकारी छिपाता है। विकल्प D डेटा की विश्वसनीयता को नष्ट करता है। परीक्षा टिप: व्यक्तिगत डेटा वाले प्रश्नों में सहमति, गोपनीयता और पारदर्शिता को प्राथमिक नैतिक सिद्धांत मानें। / Ethical data collection requires clearly informing the person, obtaining informed consent, and protecting personal privacy. Therefore, option A is correct. Option B lacks permission and transparency, while option C conceals how the data will be used. Option D destroys data reliability. Exam tip: In questions involving personal data, look for consent, privacy, and transparency as core ethical principles.
A. समस्या को उसके प्रमुख पहलुओं के साथ व्यवस्थित रूप से समझने के लिए/To understand a problem systematically through its key aspects
Explanation
Simple Explanation
समस्या कैनवास समस्या को स्पष्ट रूप से परिभाषित और समझने का उपकरण है। इसमें प्रभावित लोगों, हितधारकों, कारणों, प्रभावों और आवश्यकताओं जैसे प्रमुख पहलुओं को व्यवस्थित किया जाता है। इसलिए विकल्प A सही है। विकल्प B डेटा की सफाई से संबंधित है, जबकि विकल्प C समाधान के कार्यान्वयन से संबंधित है; ये समस्या को समझने के लिए प्रयुक्त समस्या कैनवास के कार्य नहीं हैं। / A problem canvas is a tool for clearly defining and understanding a problem. It organizes key aspects such as affected people, stakeholders, causes, impacts, and requirements. Therefore, option A is correct. Option B concerns data cleaning, while option C concerns implementing a solution; neither is the purpose of a problem canvas used for problem understanding.
कृत्रिम बुद्धिमत्ता परियोजना चक्र का पहला चरण समस्या का दायरा तय करना (Problem Scoping) है। इस चरण में समस्या, हितधारकों, आवश्यकताओं और सफलता के मानदंडों को स्पष्ट किया जाता है। डेटा का अन्वेषण, मॉडलिंग और मूल्यांकन बाद के चरण हैं, क्योंकि डेटा पर कार्य करने या मॉडल बनाने से पहले समस्या को स्पष्ट रूप से परिभाषित करना आवश्यक है। / The first stage of an artificial intelligence project cycle is Problem Scoping. This stage defines the problem, stakeholders, requirements, and success criteria. Data exploration, modelling, and model evaluation occur later because the problem must be clearly defined before working with data or building a model.
A. किसी रिकॉर्ड के एक आवश्यक डेटा-क्षेत्र में जानकारी उपलब्ध नहीं है।
Explanation
Simple Explanation
लुप्त मान दर्शाता है कि किसी रिकॉर्ड के एक या अधिक आवश्यक डेटा-क्षेत्रों में मान दर्ज नहीं है या उपलब्ध नहीं है। इसलिए विकल्प A सही है। विकल्प C में डुप्लिकेट रिकॉर्ड की समस्या बताई गई है, जबकि विकल्प D में गलत इकाई की समस्या है; ये दोनों डेटा-गुणवत्ता की अलग समस्याएँ हैं। / A missing value means that a value in one or more required data fields of a record has not been recorded or is unavailable. Therefore, option A is correct. Option C describes duplicate records, while option D describes an incorrect unit; both are different data-quality problems.
A. त्रुटियों, असंगतियों और अनुपस्थित मानों को संभालकर डेटा की गुणवत्ता में सुधार करना/To improve data quality by handling errors, inconsistencies, and missing values
Explanation
Simple Explanation
आंकड़ा सफाई में त्रुटिपूर्ण प्रविष्टियों, असंगत प्रारूपों, डुप्लिकेट रिकॉर्ड और अनुपस्थित मानों की पहचान करके उन्हें सुधारा, हटाया या उचित तरीके से संभाला जाता है। इससे डेटा की गुणवत्ता तथा विश्लेषण की विश्वसनीयता बढ़ती है, इसलिए विकल्प A सही है। विकल्प D के विपरीत, बिना जाँच के मानों को यादृच्छिक रूप से बदलना नई त्रुटियाँ उत्पन्न कर सकता है; यह डेटा सफाई नहीं है। / Data cleaning involves identifying and correcting, removing, or properly handling erroneous entries, inconsistent formats, duplicate records, and missing values. This improves data quality and makes analysis more reliable, so option A is correct. Unlike option D, randomly changing values without examination can introduce new errors and is not data cleaning.